
Training AI on Copyrighted Books: Why the Law Is Still Unsettled

The article examines whether training AI models on copyrighted books is legal, using recent court rulings and expert commentary to show the law remains unsettled.
The core issue is that copyright law, last updated in 1976, predates modern AI, leaving judges to interpret 50-year-old principles for questions that could shape the industry.
The article highlights two landmark cases that point in opposing directions. In a case against Anthropic, Judge William Alsup ordered the company to pay a $1.
5 billion settlement, but ruled that training AI models on copyrighted works was lawful.
He framed the training as analogous to a writer reading literature, distinguishing between using or consuming a work and copying it.
Attorney Cathy Gellis views this as favorable for AI companies, noting that even a large fine is comparatively small for a company projecting $200 billion in annual revenue by 2028.
In contrast, the article cites Thomson Reuters v.
Ross Intelligence, where Judge Stephanos Bibas ruled that training on copyrighted content to build a directly competing AI legal platform was not fair use.
Jason Henderson, a senior attorney, explains that courts tend to approve AI training when it does not directly compete with the source’s market, but frown upon it when it does.
Fair use law is central to these disputes, with judges weighing factors like purpose, amount used, and market impact.
The article also distinguishes between training on copyrighted works and copyrighting AI-generated content. It cites Thaler v.
Perlmutter, where a fully AI-generated work was ruled not copyrightable, raising questions about proving whether a work was AI-assisted.
Experts suggest the law will remain unresolved until later litigation settles the competing interpretations, though initial rulings are already shaping industry behavior.


