• Tiger Jerusalem@lemmy.world
    link
    fedilink
    English
    arrow-up
    65
    arrow-down
    3
    ·
    edit-2
    9 months ago

    Reddit is a trove of user built content under the guise of community. What Spez did was to say “thanks for all the free work, suckers!”, put a price sticker on it, and laughed all the way to the bank.

    And this is why I’m not active on any Internet community anymore. Nevermind, I guess I just can’t help myself…

  • garibaldi_biscuit@lemmy.world
    link
    fedilink
    English
    arrow-up
    62
    ·
    9 months ago

    This is what the 3rd party access to API was really all about.

    When API access was allowed , all reddit content was effectively free: They needed to ban 3rd party apps so they could sell the accumulated content. I expect using content to train AI also factors into it.

  • Verserk@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    45
    ·
    9 months ago

    Considering some of the very wrong and upvoted domain specific knowledge I’ve seen on Reddit over the years I’m not sure the training data is going to be useful for much beyond what every other model can do.

    • 【J】【u】【s】【t】【Z】@lemmy.world
      link
      fedilink
      English
      arrow-up
      28
      ·
      9 months ago

      The legal advice in /r/legaladvice was some of the worst garbage I’ve ever seen. I have zero doubt numerous had bad outcomes, at best wasting money and time, at worst spending years in jail because of things that sub told them to say and do. Zero doubt.

    • peopleproblems@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      9 months ago

      I can only assume they are training some specific model for something appearing more human like.

      As useless as that will be considering how fucking wildly different we type

    • SurRoulettes@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      9 months ago

      I wouldn’t be surprised if comments become their intellectual property through some terms of services bullcrap

  • Voyajer@lemmy.world
    link
    fedilink
    English
    arrow-up
    33
    arrow-down
    1
    ·
    9 months ago

    This is why I don’t blame anyone for editing/deleting their post history on reddit.

    • bcron@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      9 months ago

      It’s gonna be trained on everything, even the stuff from 2009, so I’m expecting less of that and more random ‘my fedora chortles intensify’ word salad

  • NutWrench@lemmy.world
    link
    fedilink
    English
    arrow-up
    27
    arrow-down
    2
    ·
    9 months ago

    Reddit is all bots, porn, ads and political shit posts. Good luck getting any useful training content out of that.

    • ladicius@lemmy.world
      link
      fedilink
      English
      arrow-up
      10
      arrow-down
      1
      ·
      9 months ago

      Maybe that’s the point? Training the AI to produce the blabbering bullshit that’s preferred in social media?

    • PoliticalAgitator@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      ·
      9 months ago

      They don’t care if the AI produced is useful, they just want to milk as much money from their content as they can.

      The API changes were almost certainly just the groundwork for this and I called it at the time. The ridiculous pricing model for API access is because it’s aimed at the hottest tech companies, not third party app developers.

      The enshittification continues because it’s what neoliberalism demands. They’ll sell your content and the data they have about you and still show you ads, because that’s the most profitable. Ethics and product quality don’t even enter into it.

  • ozoned@lemmy.world
    link
    fedilink
    English
    arrow-up
    21
    ·
    9 months ago

    “Reddit has given access to YOUR conversations and posts to AI companies.”. FTFY

    These were created by people, for peoole, and I will ALWAYS disagree that this data is Reddit’s or any other platforms.

    Don’t forget your direct messages aren’t end to end encrypted on Reddit, so now AI will be trained on your craziest “private” conversations

    • butterflyattack@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      9 months ago

      now AI will be trained on your craziest “private” conversations

      I have no idea what horrible thing this will do to an LLM but I’m kind of curious.

    • DocMcStuffin@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      9 months ago

      There’s one good news. Reddit didn’t want to pay to move all the old DMs to the new chat infrastructure. So they deleted them.

      • hdnsmbt@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        ·
        9 months ago

        Pretty sure they just didn’t migrate to the new data structure and didn’t actually delete the raw data. They’re effectively deleted for users but not for Reddit.

  • Bobmighty@lemmy.world
    link
    fedilink
    English
    arrow-up
    20
    ·
    9 months ago

    With reddits severe bot problem, it’ll be like training on unfiltered sewage. Garbage in, garbage out.

  • SVcrossDO@lemmy.world
    link
    fedilink
    English
    arrow-up
    17
    ·
    9 months ago

    Damn it. I haven’t deleted my account due to how many people I’ve supported and helped, I stopped using it while ago. It seems I’ll have to.

  • Yokozuna@lemmy.world
    link
    fedilink
    English
    arrow-up
    15
    ·
    9 months ago

    Good thing I scrubbed all of my posts and comments that I could. Fuck that site, straight up and down.

  • aidan@lemmy.world
    link
    fedilink
    English
    arrow-up
    12
    ·
    9 months ago

    *laughs villainously* This is all going to plan, now there will be some chatbot spewing my insane beliefs