It actually does make a difference. The genie is out of the bottle partially but it depends a lot on what we allow it it be used on. If we sit idly and allow for ingesting all what’s written for instance, including whats currently written and let bros make derivative works for a quick buck then we mostly killed the writer’s incentive to write or publish. If we slow down and not allow ripping one another off it could have the opposite effect and trully lift all the boats at once. There’s a concurent thread about Sarah Silverman suing OpenAI, that’s what Im talking about as well here
You can slow things down, but not by more than a few years, because of the gradual democratization of training foundation models. Right now training a model competitive with chatgpt can be done for $150K (microsoft orca 13b). In a few years the cost will be low enough that individuals can train models. At that point regulating it will require draconian dictatorships.
I’m also very wary of the copyright angle on this, because just like we don’t prohibit people from learning copyrighted materials in their brains, it feels very wrong to regulate how we train digital brains. I’m ok with forbidding the output of copies of individual existing copyrighted works, but we already have laws on the books for that. I find it downright immoral to prohibit the generation of works “in the style of”. That again reeks like the kind of draconian society I don’t want to live in.
People will always be willing to pay for human-made art, just like we pay more for handmade pots, even though machines can make them better, so I think the doomsayers who predict the end of art are flat out wrong. Easy access to mass-generated AI content could be the best thing that happened to true artists, just like chess AI that can beat every human player was the best thing that happened to the chess world. We need labeling laws that show the origin of works so people can choose whether they want artificial or human-made, but please not another extension of the copyright regime to be even more suffocating and hostile of cultural flourishing.
>> You can slow things down, but not by more than a few years, because of the gradual democratization of training foundation models.
Just to be clear, what's being "democratised" is the fine-tuning of second-tier, inferior-performance models; or pre-training of third-tier ones. In the game of training large neural nets, the players that can afford to train the largest models with the most amount of data and compute at any given time will continue to dominate for the foreseeable future.
To make it plain, maybe in a couple of years you'll be able to train GPT-4 on your student laptop (unlikely, but let's allow it for the sake of argument). You'll still not be able to get anywhere near the performance of GPT-6 or whatever OpenAI and Google will be able to train by then.
Academics, hobbyists and smaller companies will continue to play second fiddle to large corporations as long as the dominant paradigm is more data and more compute.
Capabilities of the open-source models are only increasing over time by objective measurement. Yes, every one of them is demonstrably inferior to GPT-4, but we have historical precedent that the cost of compute only ever goes down.
Additionally, assuming the leaked details given here are accurate, there might not be a GPT-6. This entire approach of AI via language models very well could be approaching a local maximum and/or have already reached the point of diminishing returns.
If that is the case, OpenAI's moat is guaranteed to run dry. It should be telling that very few of the improvements over the past few months involve the base model, rather they are value-adds like plug-ins and hooking it up to a VM, things that are not protected by training difficulty.
I don't think there's a problem with the pace of "progress", the problem is what it is feed and we could definitely make changes around that. For example, if I were an author I wouldn't want this stuff to be be feed in without my permission.
> If we sit idly and allow for ingesting all what’s written for instance
It's already happened. What do you do now? For decades, it's happened regardless of the robots.txt so you have to assume it's all been ingested by it all, everywhere. What now?