The legality of training artificial intelligence models on copyrighted books remains unresolved, as courts grapple with applying decades-old copyright law to rapidly evolving technology. Recent rulings have offered partial clarity, but the core question—whether AI training constitutes fair use—continues to divide judges and legal experts.
The Legal Landscape: Fair Use and AI Training
At the heart of the debate is the doctrine of fair use, which permits limited use of copyrighted material without permission for purposes such as criticism, comment, news reporting, teaching, or research. Courts weigh four factors: the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the potential market. In AI training, the key issue is whether using copyrighted works to train a model is “transformative”—that is, whether it adds new expression or meaning rather than merely copying the original.
According to Cathy Gellis, an attorney specializing in intellectual property and technology, “Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work, consuming the work, reading the work.” This distinction is central to how some judges view AI training: as akin to a human reading a book to learn, rather than copying it verbatim.
Recent Court Rulings: Mixed Signals
In a landmark case last year, Judge William Alsup ordered Anthropic to pay $1.5 billion to a group of authors whose works were used to train its AI models. However, the judge also ruled that Anthropic’s training itself was lawful; the penalty was for using pirated copies from shadow libraries. This nuanced decision highlights the complexity: the act of training may be permissible, but the source of the data can create liability.
In contrast, a separate case involving Thomson Reuters and Ross Intelligence produced a different outcome. Judge Stephanos Bibas ruled that Ross’s use of Reuters’ content to build a competing legal research platform was not transformative and thus not fair use. The key difference: Ross’s product directly competed with the original, whereas AI models like Anthropic’s are not designed to replace the books they train on.
Why the Rulings Differ
Legal experts point to the competitive nature of the use as a decisive factor. Jason Henderson, Senior Attorney at JWL International, explains, “What’s tending to win is if what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it… If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay.” This distinction could shape future cases, as authors may argue that AI-generated content competes with their own works, though that argument has yet to prevail.
Implications for Authors and AI Companies
For authors, the uncertainty is troubling. Many feel their livelihoods are threatened by AI tools trained on their work without consent or compensation. Yet the legal system has not provided a clear answer. The Copyright Act of 1976 predates the internet, let alone artificial intelligence, leaving judges to interpret outdated guidelines for novel situations.
For AI companies, the rulings offer some reassurance. The $1.5 billion fine against Anthropic, while significant, is a fraction of the company’s projected revenue, suggesting that even adverse outcomes may not deter AI development. Gellis notes that the decision is “generally good news for AI training” because it treats AI learning as analogous to human reading.
The Road Ahead: Pending Litigation and Unanswered Questions
Many AI companies remain embroiled in litigation, and the legal landscape is far from settled. “What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things,” Gellis said. “It’ll take later states of litigation to figure out which one will prevail.” Until then, the rules governing AI training on copyrighted works will remain in flux.
Conclusion
The question of whether AI can legally train on copyrighted books is far from resolved. Recent rulings have provided some guidance, but they are inconsistent and subject to appeal. For now, both authors and AI companies must navigate a legal gray area that could shape the future of creativity and technology.
FAQs
Q1: Is training AI on copyrighted books considered fair use?
Not definitively. Courts have issued conflicting rulings. In some cases, training has been deemed lawful if the use is transformative and doesn’t directly compete with the original. In others, it has been ruled infringement, especially when the AI product competes with the source material.
Q2: What was the outcome of the Anthropic case?
Judge William Alsup ordered Anthropic to pay $1.5 billion for using pirated books from shadow libraries, but he also ruled that the AI training itself was lawful. The penalty was for copyright infringement in sourcing the data, not for the act of training.
Q3: How does the Thomson Reuters v. Ross Intelligence case affect AI training?
In that case, the court ruled that Ross’s use of Reuters’ content to build a competing legal platform was not fair use. This suggests that if AI training results in a product that directly competes with the original work, it is more likely to be considered infringement.
Disclaimer: The information provided is not trading advice, Bitcoinworld.co.in holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

