• schipelblorp@sh.itjust.works
    link
    fedilink
    arrow-up
    23
    ·
    edit-2
    1 day ago

    File under obvious?

    This is likely a desperate attempt to avoid model collapse, where AI feeds itself data and starts hallucinating more and more often. But being able to identify AI data means they can keep AI content out of the data and focus on content created by “the agents that live in reality”, aka the entities formerly known as human beings.

    Of course that means we also can start filtering AI content out of our feeds and search engines and web sites and training data. That makes watermarked AI output exactly as valuable as it is actually is to society (<0), so I doubt Anthropic would make this mandatory. And if it did make it mandatory, there are plenty of models that aren’t doing this that people can use.

    This is just the pointless thrashing about of a drown victim, as far as I can tell.

    • sanpo@sopuli.xyz
      link
      fedilink
      arrow-up
      9
      ·
      1 day ago

      Or it’s a compliance thing.

      EU now requires slop disclosure, even text in certain scenarios.

  • brucethemoose@lemmy.world
    link
    fedilink
    arrow-up
    13
    ·
    edit-2
    1 day ago

    I like 404, but to be blunt, the author doesn’t understand how absolutely tiny the “nudge” is.

    Nor how LLM sampling works.

    It’s not figuratively imperceptible; it’s literally below the noise floor of default sampling parameters, unless you run Claude at ~zero temperature, which no one does for this kind of stuff.

    This kind of token bias only becomes statistically significant in larger bodies of text. Word to word, it does basically nothing.


    If they have a problem with imprecise word choice, as they do in the article… Well, yes. Thats the issue with LLM sampling. It’s the elephant in the room.

    To me, basic, temperature-based sampling with top-k/top-p was a “bandaid” to fix weird self-feedback loops with autoregressive research artifacts, like looping and repetition. It was a hack. And they just… commercialized it and never fixed it.

    There are tons of interesting papers on alternatives to sampling. There tons of interesting implemented improvements (I’m partial to sigma-n/adaptive-p, tuned token bias, and constrained output grammar), but of course Claude offers none of these, and no one can do anything about it because it’s a closed model system.

  • imsure@lemmy.world
    link
    fedilink
    arrow-up
    11
    ·
    edit-2
    1 day ago

    Maybe I’m missing something, but I don’t get this “oh they don’t care about the wording” thing. If you use AI you get what is generated by a neuronet that is not perfect in the first place. It’s not like someone deprives you of you personality, it is not your text anyway. If you want quality and vibrant text, you have to write it yourself. So I don’t see what’s the point in moaning over this “gray vs overcast” thing

  • isleepinahammock@lemmy.blahaj.zone
    link
    fedilink
    English
    arrow-up
    7
    ·
    1 day ago

    I honestly don’t see anything wrong with this. AI text generators are only good for writing content no one cares about. If no one is meant to care about it, then the subtle difference between synonyms really doesn’t matter. This is like someone putting a watermark inside the weaving pattern of an industrial knitting loom, and then someone complaining that this invalidates the artistic expression of knitting. It’s not the watermark degrading the artistry, but the automated knitting loom.

    Al writing is already soulless and devoid of artistic merit. Why complain about making it more identifiable? I really do not understand the objection at all.