Courts Apply Fair Use to AI Training on Copyrighted Books
The legality of training artificial intelligence models on copyrighted literary works remains a deeply contested legal frontier, with recent judicial decisions beginning to carve out a precedent that largely favors AI developers. Last year, Judge William Alsup ordered Anthropic to pay a $1.5 billion copyright settlement to a coalition of authors whose works were incorporated into the company’s language models. Despite the steep financial penalty, the ruling explicitly deemed the training process itself lawful. Judge Alsup distinguished between unauthorized reproduction and the act of reading, noting that large language models ingest text to analyze patterns and generate novel outputs, a process he compared to an author studying existing literature. Legal experts, including intellectual property attorney Cathy Gellis, view the decision as a significant win for the AI industry. She emphasized that traditional copyright law centers on the prohibition of copying rather than the consumption or application of creative works. The broader legal environment remains fragmented, as the current U.S. Copyright Act has not been substantively updated since 1976. Judges are now tasked with applying mid-twentieth-century statutes to rapidly evolving generative technologies. Central to these disputes is the fair use doctrine, which permits limited use of protected material without permission if the application is deemed transformative. Courts evaluate several factors, including the purpose of the use, the quantity of material utilized, and the potential impact on the original work’s market. Jason Henderson, a senior attorney at JWL International, noted that judicial outcomes frequently hinge on competitive overlap. He pointed to the recent Thomson Reuters v. Ross Intelligence case, where a court ruled against fair use after determining that the defendant’s AI platform directly competed with the plaintiff’s proprietary legal database. Conversely, when AI training does not directly displace the original creator’s market, courts have been more inclined to find permissible use. The judicial distinction between training data and generated output is further complicated by separate rulings on AI authorship. In Thaler v. Perlmutter, a federal court established that works created entirely by AI lack copyright protection, raising complex evidentiary questions regarding human-AI collaboration in creative processes. These overlapping legal battles mean the industry operates in a state of prolonged uncertainty. While initial verdicts are already influencing corporate strategy and development pipelines, subsequent litigation across multiple jurisdictions may ultimately redefine the boundaries of permissible data scraping. For now, AI developers must navigate a patchwork of preliminary rulings that acknowledge the transformative nature of model training while carefully avoiding direct market displacement of established copyright holders.
