Unsealed Briefs Show Execs Knew AI Book Piracy Was Illegal
Unsealed court filings in the ongoing class action Alter v. OpenAI and Microsoft have intensified scrutiny over the developers of leading large language models, revealing alleged corporate awareness that training AI on pirated literature violated copyright law and threatened author livelihoods. The documents, released September 21 in Manhattan federal court, form part of a lawsuit filed by the Authors Guild and a coalition of prominent writers. According to the plaintiffs briefs, OpenAI intentionally incorporated millions of copyrighted books hosted on Library Genesis, a widely known piracy site, to train its GPT models. Internal communications indicate Microsoft executives were informed of this data sourcing as early as April 2019. Rather than addressing copyright concerns, company leadership reportedly weighed legal exposure against public relations risks. OpenAI executives later initiated a cleanup operation dubbed Project Clear in summer 2022, directing the deletion of Library Genesis references from internal systems after concerns surfaced about potential media scrutiny. The filings underscore a pattern of internal prioritization of model development over creator rights. Executives and researchers repeatedly acknowledged that advanced generative systems would substitute human creative labor. In May 2020, OpenAI policy director Jack Clark cautioned that improved language models would increasingly displace genre fiction writers, predicting the company would proceed despite artist pushback. By 2022, OpenAI research lead Tarun Gogineni outlined goals to autocompleting stalled literary series, dismissing author complaints about unauthorized dataset usage as acceptable economic disruption. Internal memos also suggested that machine-generated content would eventually be consumed by other machines, raising concerns about cultural degradation. Authors Guild leadership characterized the disclosures as proof of deliberate copyright infringement and a systemic devaluation of professional writing. The plaintiffs argue that the unchecked training on stolen literary works poses an existential threat to publishing and American cultural production. The case, led by counsel Justin A. Nelson of Susman Godfrey, has reached a critical procedural stage following the motion for partial summary judgment. A court hearing is scheduled for early 2027, with both sides expected to submit further briefing in the coming months. As artificial intelligence capabilities accelerate, the litigation highlights mounting legal and economic pressures on creative industries grappling with unauthorized data use and workforce displacement.
