• monobot ( monobot@lemmy.ml ) 
        link
        fedilink
        arrow-up
        45
        ·
        3 years ago

        I just checked and my account on lemmy is 3 years old, I was waiting for all of you. At moments lemmy looked like it will take of by it self, but last few months was pretty quiet. I hope this is the push this community needs to succeed.

        Reddit got too big for my taste since digg joined in, I don’t think this will kill it ( they have the data how many people is using it outiside official apps), but let’s make space out of it.

        • All it took was for a few power users, aka content thieves, someone looks MrBabyMan who’s spent all conscious time submitting entertaining content from other less user friendly platforms, to convince a majority of users on the internet to pay attention to reddit more.

          Simply getting more content, regardless of its quality or dubious origin, only as long as it is interesting or entertaining, will draw more users here. This is probably not what people want, but that’s what it will take for Lemmy to take off.

        • ture ( ture@rational-racoon.de ) 
          link
          fedilink
          arrow-up
          7
          ·
          3 years ago

          If one was still browsing reddit using rif and seldomly old reddit it somehow still felt similar to what it was 10 or more years ago. Over time I ditched Facebook and some other social media I used but Reddit somehow stayed. Maybe because it was the “anonymous” one, the one I just used for myself without sharing my account with my real-life friends.

          Anyway thanks for waiting for us, took some time for me to get up and leave reddit. I hope others will follow, but so far even for me as a software dev/ architect it was quite a change to switch to fediverse services. Maybe it’ll be smoother when you’re joining bigger servers but let’s see what the future brings. As a fan of Foss I’d really like to see this thing grow.

  • pancake ( pancake@lemmy.ml ) 
    link
    fedilink
    arrow-up
    69
    ·
    3 years ago

    Do you expect this is a reflection of how Reddit will handle relations with its investors?

    Holy shit, they killed him right there. They have put the thread in “sort by new” mode and I bet it’s just to bury that bomb as deep as they can.

  • I feel like this move has nothing to do with investors and everything to do with setting the standard for big corps like Microsoft and Google to be able to scrape their massive amount of data to train next gen AIs. They know they have HUGE amount of data from now and for years and years ago. Content, created by others, then sold for enormous profit.

    • Nina ( misnina@lemmy.ml ) 
      link
      fedilink
      arrow-up
      27
      ·
      3 years ago

      I mean AI is already stealing all art and images on the web without paying anything. They could just literally scrape and pay nothing. Web scraping isn’t illegal, they already do it, why would they pay anyone? Unless the law catches up about the rights to manufacture AI content based on ill-gotten data, then why would they pay what they don’t have to?

      • Its not even considered stealing to make training data, you can disagree, but that has to be Maschine readable by law, everyone knows that its Algorithms scraping the web for data, if they see a image that says don’t scrape they don’t take it but most images don’t have such a attachment.

        • Nina ( misnina@lemmy.ml ) 
          link
          fedilink
          arrow-up
          1
          ·
          3 years ago

          Could you please point me to legal definitions, in court or otherwise, that say it is not violating my copyright license to directly use my artwork in any shape or form for a non-fair use product? As in, a service you pay money for to create things based on the training data it has taken from me, is not fair use. Or point me to the legal definitions where I lose my copyright by posting things online? Allowing to scrape is not the same thing as giving derivative copyright license permissions. You aren’t disagreeing with me, you’re disagreeing with my legal rights.

            • hmm. That talks of data mining not of derivative work from the mined data. … I’ll leave the discussion to others, though.
              What effect would a Creative Commons “ND” clause have? Is it reserving the right to make derivatives? And would machine generated stuff even count as a derivative?

              DuckDuckGo translate:

              Copyright and Related Rights Act (Copyright Act) § 44b Text and Data Mining
              (1) Text and data mining is the automated analysis of one or more digital or digitized works in order to obtain information, in particular about patterns, trends and correlations.
              (2) Reproduction of lawfully accessible works for text and data mining is permitted. The reproductions must be deleted when they are no longer required for text and data mining.
              (3) Uses pursuant to subsection (2) sentence 1 shall only be permitted if the rights holder has not reserved them. A reservation of use for works accessible online is only effective if it is made in machine-readable form.

            • Nina ( misnina@lemmy.ml ) 
              link
              fedilink
              arrow-up
              2
              ·
              edit-2
              3 years ago

              Translation: (1) Text and data mining is the automated analysis of one or more digital or digitized works in order to obtain information, in particular about patterns, trends and correlations.

              (2) Reproduction of lawfully accessible works for text and data mining is permitted. The reproductions must be deleted when they are no longer required for text and data mining.

              (3) Uses pursuant to subsection (2) sentence 1 shall only be permitted if the rights holder has not reserved them. A reservation of use for works accessible online is only effective if it is made in machine-readable form.

              None of that says anything about creating profitable derivative work. In fact is specifies patterns, trends and correlations, which does not lead me to believe it is protecting visual works created from this data, those kind of things are only used to inform things, like information science.

              • Algorithms trained with that training data are considered to make new works, so your only change to stop your work getting into a Algorithm is by going against the ones that make the training data. And they will always pull the science card.

      • What do you mean by stealing? The data remains, all they do is learning from something which is public

        What is different to Googles approach, they are just watching and learning Why is it treated so differently when it essentially does nothing new, but uses the data in a different way

        • Nina ( misnina@lemmy.ml ) 
          link
          fedilink
          arrow-up
          4
          ·
          3 years ago

          they are just watching and learning Why is it treated so differently

          Because it isn’t human. It isn’t watching and learning, it is being fed my creative content as data that I have not allowed nor have been compensated for, which is then turned around and sold as a service. My work is being consumed for commercial uses by an inhuman who does not have fair use education rights, with the sole intent to create a profitable product, and I’m getting nothing. I have legal rights, no matter where I post my work, to retain my copyrights and I have the right to not consent to improper use of my works that do not align with the licenses I have chosen to give it. Websites ask for a licenses in their ToS to be able to even just display and share my artwork when I upload it. When I create an image, I am given ownership of it’s copyright to control the use, distribution, and right to create derivatives. This isn’t a fuzzy area, it’s very clear. If an artist did not consent to their artwork being used as training data for a non-fair use reason, it is stealing their works.

          And no, it’s not fair use under education. Copyright exists for human protection and uses. It isn’t being used for ‘learning’ it’s used as data to be repackaged and sold. Google’s use of it showing up in search is to link back to posts that contain my work, retain my copyright, and are not derivatives. If you mean by captchas, yeah capchas are pretty bullshit.

          And circling back to my original post. So? AI companies aren’t paying for their image training data, so why would they pay for reddit’s api?

          • slacktoid ( slacktoid@lemmy.ml ) 
            link
            fedilink
            arrow-up
            2
            ·
            3 years ago

            I feel the bigger problem with these AIs is more how they are solely being used to improve profits and productivity, these only affect the capital owners. None of that is going to improve the laborer (i.e., the artist, the coder, the writer, the people who create value from capital). This is only going to get worse. We are being normalized to automation and AI with the use of self-checkout.

            Also, about Reddit training data, I think they are too late to the party. The weights they were needed for are made. I do not think they are the exclusive source of specialized information, and (I hope) they are going to find out. They are just going to further show how silly the free market and the stock market are. The people who require the data will probably have other ways of getting it. r/datahoarders and people like that come to mind. Reddit is only making new data hard to access which, which they are not (and hopefully never) an exclusive source of.

            • Nina ( misnina@lemmy.ml ) 
              link
              fedilink
              arrow-up
              3
              ·
              edit-2
              3 years ago

              Yeah, AI can totally exist and be useful, but currently it’s in the hands of tech dudes and admins who have a terrible track record with developing things responsibly and over hyping and masking flaws. It’s used to make a profit at the colossal detriment to humans. It’s used to hurt us currently, not help at all.

              I think the training data from reddit probably only used the API because it was easier and free. And if no longer free, there’s nothing pointing to them actually paying for it. It’s not like reddit is the only data, they very much likely already have web scrapers for other uses that they can just tune for reddit.

    • The thing I worry about whenever someone mentions this angle: What about Lemmy content? As the community moves away from the commercial platforms in favor of Lemmy, Bluesky, Mastodon etc. Then does that lower the legal barrier for AI companies to train on all this content for free? Is that shift in the legal vulnerability of public content something that users consider? Is that desirable to most users? Are people thinking about that?

    • That’s an interesting but definitely plausible take on the whole thing. 12000$ for 50mio requests is B2B pricing. For a company like Openai/Microsoft that’s not even worth thinking about if you get all of that precious training data for it…

        • argv_minus_one ( argv_minus_one@beehaw.org ) Banned
          link
          fedilink
          arrow-up
          1
          ·
          edit-2
          3 years ago

          Right, but these are big companies with lots of talented programmers on hand. If anyone can overcome such an obstacle, it’s them.

          Also, Google and Microsoft already have a search index full of Reddit content to scrape.

          • You are right. You would need a team of skilled scrapers and network engineers though would know how to get around rate limiters with some kind of external load balancer or something along those lines.

            • Rate limiters work on IP source. This is easily bypassed with a rotating proxy. There are even SaaS that offer this. The trick is to not use large subnets that can be easily blocked. You have to use a lot of random /32 IPs to be effective.

              • Could they be doing that already because of the still open API of Reddit and that will soon change? I just feel like it’s easier for them currently and it will be tougher once the API changes are implemented.

                • No. Search engines fetch pages using plain old HTTP GET requests, same as how browsers fetch pages. There is some difficulty in parsing the HTML and extracting meaningful content, but it’s too late: the HTML is already stored on Google/Microsoft servers, ready for extraction, and there’s nothing Reddit can do to stop them.

                  Reddit can make future content harder to extract, but not without also making it invisible to search engines, which would cause Reddit to disappear from Google Search and Bing. That would destroy Reddit even faster than a mass moderator exodus will.

                  That’s why I say trying to charge money for AI training data is a fool’s errand. These facts make it impossible. That doesn’t mean Spez won’t try, but it does mean he won’t succeed.

    • DrQuint ( DrQuint@lemmy.ml ) 
      link
      fedilink
      arrow-up
      4
      ·
      3 years ago

      Interesting. I wonder if they already got an offer that matches their new API pricing, and they decided to up everything to match that cost and avoid being sued later.

      Like, there seems to be some urgency between them announcing and upping the price. What was it? Is this the reason? A confirmed, extremely wealthy and extremly naive buyer?

  • hydra ( hydra@lemmy.one ) 
    link
    fedilink
    arrow-up
    34
    ·
    3 years ago

    The enshittification of Reddit has been evident ever since the new design rolled out. Unusable on mobile devices. Does less, using more memory and bandwidth.

    • When “New” Reddit came out, it was just shockingly bad. If they didn’t keep old.reddit.com online, they would have killed the site then. Until very recently I couldn’t even view all child comments within the main thread, and it still takes at least twice as long to load any page.

      Coming to Lemmy has been a breath of fresh air. The site is much more responsive than Reddit despite most instances running on a single VPS or something.

      • hydra ( hydra@lemmy.one ) 
        link
        fedilink
        arrow-up
        11
        ·
        edit-2
        3 years ago

        That’s because the Lemmy webapp focuses on being lean and functional rather than shoving as much telemetry and megabytes of JavaScript bloat as they can to do LESS than the old Reddit webapp could while using 10-20 times MORE resources.

        New Reddit is not completely unusable on an intermittent, crammed full 3G connection where Old Reddit just works, but is known to be actively user hostile and somehow cramming full a huge 1080p phone display with only 2 or 3 comments and having to preload for hours for a full thread.

        Lemmy server is also blazing fast and being written in Rust which encourages memory safety does help it function better on smaller instances and serve both local and federated clients faster despite having less resources to do so. I really really hope this replaces Reddit in the mainstream and people learn basic concepts about federated media to future proof the free Internet.

        It has been a real breath of fresh air and so far it seems more sustainable to spread the bandwidth between smaller instances than let a megacorp fund the infrastructure to serve everyone in a walled garden which will later be enshittified into garbage once a critical mass is already lured in.

        EDIT: look for yourself and notice the difference on Firefox’s dev console network tab

        • You’re right! The front page of Reddit is nearly 8x larger than Lemmy.ml, and took almost 7x longer to load than Lemmy.

          Uncached loading results:

          Lemmy: 3.3 MB, 39 requests in 1.85 seconds
          Old Reddit: 6.3 MB, 60 requests in 4.53 seconds
          New Reddit: 24.5 MB, 351 requests in 12.21 seconds

  • I was reading an article from The Verge - they’ve linked all the juicy parts of the AMA. I don’t have a Reddit account anymore since hearing about what their plans were a few weeks ago. It looks like Reddit is facing a very real, self inflicted meltdown.

    Fyi - it looks like the reviews for Reddit on the Google Play Store have been locked as well to prevent review bombing.

  • Burn it to the ground! I wish all top subreddits had the balls to go dark indefinitely to the point they have to backpedal or forcibly take over the subreddits. Burn it to the fucking ground!

  • anilm ( anilm@kbin.social ) 
    link
    fedilink
    arrow-up
    12
    ·
    3 years ago

    I’m surprised they didn’t replace /u/spez with woman; then have her make this announcement. Then he could triumphantly return and claim they will do better from now on. … you know, like the last time they made unpopular changes.

    • yep, I had a similar train of thought

      [spez] screwed up so bad, and it’s obvious he’s quite stupid…I’m wondering if he was given the Ellen Poa treatment. The investor’s knew what changes they wanted to implement to increase revenue, had him implement it, then made him take all the flak. Next, they’ll replace spez with someone else, seemingly roll back a few of the changes, and say they’re working with the reddit community. The thing is that it’s obvious that reddit is irreparably infected. The diseases cannot be eradicated because the disease is corporate greed. Reddit began its time in hospice the moment they shared they were going public. It’s over. Done. Say you final goodbyes. RIP in pieces.