▲ 49 r/webdev
Where did AI companies get legal permission to train on copyrighted data?
This is a question I have thought about a lot lately. If it wasn’t for the repositories hosted in Github and other platforms, none of the models would exist now. Specifically, I don’t remember agreeing on anything that said something about training AI models with my data when I started hosting code on Github, more than a decade ago. Yet it’s a wide known fact that the data hosted there has been used for training for a long time.
These models are used commercially and may produce substantial fragments of copyrighted code because internally, they still contain copyrighted data.
Even if the training is done on only MIT licensed code, then upon use, the original author’s name must still be reproducible.
Can someone explain this to me?
u/OkShip110 — 5 hours ago