The Irony of 'You Wouldn't Download a Car' Making a Comeback in AI Debates

FatCat@lemmy.world · 12 days ago

The Irony of 'You Wouldn't Download a Car' Making a Comeback in AI Debates

mriormro@lemmy.world · 12 days ago

You know, those obsessed with pushing AI would do a lot better if they dropped the patronizing tone in every single one of their comments defending them.

It’s always fun reading “but you just don’t understand”.

FatCrab@lemmy.one · 12 days ago

On the other hand, it’s hard to have a serious discussion with people who insist that building a LLM or diffusion model amounts to copying pieces of material into an obfuscated database. And then having to deal with the typical reply after explanation is attempted of “that isn’t the point!” but without any elaboration strongly implies to me that some people just want to be pissy and don’t want to hear how they may have been manipulated into taking a pro-corporate, hyper-capitalist position on something.

mriormro@lemmy.world · 12 days ago

I love that the collectivist ideal of sharing all that we’ve created for the betterment of humanity is being twisted into this disgusting display of corporate greed and overreach. OpenAI doesn’t need shit. They don’t have an inherent right to exist but must constantly make the case for it’s existence.

The bottom line is that if corporations need data that they themselves cannot create in order to build and sell a service then they must pay for it. One way or another.

I see this all as parallels with how aquifers and water rights have been handled and I’d argue we’ve fucked that up as well.

VoterFrog@lemmy.world · edit-2 12 days ago

They do, though. They purchase data sets from people with licenses, use open source data sets, and/or scrape publicly available data themselves. Worst case they could download pirated data sets, but that’s copyright infringement committed by the entity distributing the data without the legal authority.

Beyond that, copyright doesn’t protect the work from being used to create something else, as long as you’re not distributing significant portions of it. Movie and book reviewers won that legal battle long ago.

FatCrab@lemmy.one · 12 days ago

Training data IS a massive industry already. You don’t see it because you probably don’t work in a field directly dealing with it. I work in medtech and millions and millions of dollars are spent acquiring training data every year. Should some new unique IP right be found on using otherwise legally rendered data to train AI, it is almost certainly going to be contracted away to hosting platforms via totally sound ToS and then further monetized such that only large and we’ll funded corporate entities can utilize it.

yamanii@lemmy.world · 12 days ago

I don’t get your comment, are the pro corporate for AI or against it?

FatCrab@lemmy.one · 12 days ago

I have no personal interest in the matter, tbh. But I want people to actually understand what they’re advocating for and what the downstream effects would inevitably be. Model training is not inherently infringing activity under current IP law. It just isn’t. Neither the law, legislative or judicial, nor the actual engineering and operations of these current models support at all a finding of infringement. Effectively, this means that new legislation needs to be made to handle the issue. Most are effectively advocating for an entirely new IP right in the form of a “right to learn from” which further assetizes ideas and intangibles such that we get further shuffled into endstage capitalism, which most advocates are also presumably against.

yamanii@lemmy.world · 12 days ago

I’m pretty sure most people are just mad that this is basically “rules for thee but not for me”, why should a company be free to pirate but I can’t? Case in point is the internet archive losing their case against a publisher. That’s the crux of the issue.

FatCrab@lemmy.one · 11 days ago

I get that that’s how it feels given how it’s being reported, but the reality is that due to the way this sort of ML works, what internet archive does and what an arbitrary GPT does are completely different, with the former being an explicit and straightforward copy relying on Fair Use defense and the latter being the industrialized version of intensive note taking into a notebook full of such notes while reading a book. That the outputs of such models are totally devoid of IP protections actually makes a pretty big difference imo in their usefulness to the entities we’re most concerned about, but that certainly doesn’t address the economic dilemma of putting an entire sector of labor at risk in narrow areas.