Can AI Companies Train on Copyrighted Work?
Why authors, artists, publishers, and AI companies are fighting over training data.
A short answer
It is unsettled. AI models are trained on very large collections of text, images, audio, and code, much of it copyrighted.
AI developers generally argue this is fair use — a legal doctrine allowing limited unlicensed use, especially when the use is "transformative." Authors, artists, news organizations, and record labels argue that copying their work to build a commercial product that competes with them is not fair use.
Where the fight is happening
Most of it is in court. Dozens of U.S. lawsuits have been filed against major AI developers, and rulings so far have been mixed and fact-specific — turning on how the material was obtained, whether the model reproduces recognizable portions of it, and what market harm is shown.
Some companies have instead signed licensing deals with publishers and stock-media libraries, which creates a parallel commercial answer while the legal one is pending.
The policy debate
Options being discussed include a licensing or collective-payment system, mandatory disclosure of training data sources, opt-out mechanisms for rights holders, and leaving the question to the courts.
The U.S. Copyright Office has issued reports on AI and copyright, including on training data, but Congress has not passed legislation resolving the question.
Further reading
Last updated September 2026