Microsoft exec calls AI scraping 'largest theft of labor' in court filings
New unredacted court filings in the three-year-old copyright lawsuit between The New York Times and OpenAI and Microsoft have surfaced, revealing extensive internal admissions regarding artificial intelligence training practices and their economic impact on traditional publishing. The newly unsealed materials detail how the technology firms allegedly bypassed paywalls undetected, stripped copyright notices before model ingestion, and accumulated millions of copies of journalistic content for generative training. Internal documents from early 2024 describe the mass scraping operations as potentially threatening the economic foundations of the content supply chain. Microsoft’s director of Applied Science characterized the practice as unprecedented theft, warning that the technology could severely disrupt employment within the very sector supplying the training data. Dataset revelations show OpenAI mid-training data alone contained over ninety-one thousand copies from major news outlets, while a common crawl-derived repository held more than two million documents solely from The New York Times. Microsoft analytics indicate its Copilot answer engine reduced click-through traffic to The New York Times by as much as ninety-three percent compared to traditional search. Internal communications warn this traffic diversion could trigger a destructive cycle harming both model performance and the broader web ecosystem. Leadership statements further complicate the defendants position. Under oath, Microsoft chief executive Satya Nadella acknowledged that paywalled material should be licensed for training, stating he would have mandated OpenAI retrain its models had he known paywalls were being circumvented. OpenAI executives echoed concerns about market displacement, with internal communications noting that conversational AI directly substitutes access to original sources and poses an existential threat to journalism. The filings also document a coordinated strategy among OpenAI researchers to develop workarounds for paywall detection, which company leadership acknowledged in written correspondence. These revelations directly challenge the defendants fair use defense, particularly the legal requirement that AI training must not harm the market for original works. While federal judges have historically leaned toward fair use and the Trump administration recently filed a brief supporting unlicensed training, these newly unsealed internal records underscore the direct competitive overlap between generative models and publisher content. The companies have not responded to requests for comment regarding the unredacted materials.
