Business
Microsoft Internal Criticism Surfaces on AI Training Data as NYT Lawsuit Looms
Internal communications from Microsoft, now public due to The New York Times' lawsuit against OpenAI and Microsoft, reveal significant internal dissent regarding the methods used for training artificial intelligence models. These candid remarks describe the unauthorized scraping of millions of copyrighted texts as “astonishing theft of unprecedented proportions” and potentially the “largest theft of labor in human history.” The statements, attributed to Brent Hecht, head of Applied Science at Microsoft, underscore a deep-seated concern within the AI industry about the ethical and legal implications of its data acquisition practices.
Internal Doubts on 'Fair Use'
The revelations come to light amidst a high-profile legal battle initiated by The New York Times in late 2023. The newspaper accuses OpenAI and Microsoft of infringing copyright by using millions of its articles without permission to train AI models like ChatGPT. The lawsuit argues that these AI systems threaten the publisher's subscription, licensing, and advertising revenues. Crucially, the internal Microsoft documents suggest that the company's public argument, asserting that training AI on news texts falls under the US legal exception of “Fair Use,” is not genuinely believed internally. Hecht is quoted as admitting that if the company were to prevail in court with this argument, it “would be a complete mockery of the idea of ’fair use’.”
The 'Doom Loop' of AI Content Strategy
Further internal criticism points to a self-defeating cycle created by the AI's capabilities. Microsoft has observed that when its AI is queried about news topics, users click on links to the original newspaper articles far less frequently. This is because the AI's primary advantage is its ability to provide direct answers rather than a list of sources. This shift directly impacts publishers' ability to generate traffic and revenue. Microsoft's senior legal counsel acknowledged under oath that chatbot conversations would indeed replace visits to the original sources. The company's internal assessment warns that its AI content strategy has initiated a “doom loop” that could negatively affect the performance of its models and the broader web ecosystem.
Existential Threat to Publishers
The internal admissions highlight a stark awareness of the existential threat AI poses to the publishing industry. Nick Turley, identified as the head of ChatGPT at OpenAI, reportedly admitted internally that AI serves as “largely substitutive [for press texts], period.” He further elaborated that as AI models improve, they will increasingly replace the products offered by publishers. This admission from a key figure at OpenAI directly contradicts the public narrative often presented by AI companies. The situation is further complicated by the fact that AI models themselves rely on the very content they are threatening to displace, creating a paradoxical economic dependency.
OpenAI's Pragmatic Approach to Paywalls
Beyond Microsoft's internal reflections, quotes from OpenAI co-founder Greg Brockman offer a glimpse into the company's pragmatic, and perhaps opportunistic, approach to accessing content. When informed that a method had been found to bypass The New York Times' paywall, Brockman's reported response was simply, “oh nice.” This casual remark, juxtaposed with his elsewhere stated motivation by the “gazillions of dollars” ChatGPT could generate, suggests a focus on growth and revenue that may have overshadowed ethical considerations regarding content acquisition. These internal sentiments, now public, provide significant ammunition for The New York Times in its landmark legal challenge.
Precedent-Setting Legal Battle
The case is widely considered a critical precedent for determining the legality and ethics of AI training data practices. The New York Times' lawsuit, filed under File No.: 1:25-md-03143-SHS-OTW, seeks to establish whether the current methods employed by AI developers are protected under “Fair Use” or constitute widespread copyright infringement. The internal communications from Microsoft and OpenAI, now part of the public record, appear to validate many of The New York Times' core arguments. They suggest that the companies are aware of the potential legal and economic ramifications of their data-gathering strategies, even as they publicly defend their practices.
Broader Implications for the Web
The implications of this internal criticism extend beyond the immediate legal dispute. Microsoft's acknowledgment of a potential "doom loop" for the web raises fundamental questions about the long-term sustainability of online content creation and consumption. If AI models directly answer user queries without directing traffic to original sources, the economic model that supports journalism and other forms of content creation could collapse. This could lead to a less diverse and less reliable information ecosystem, as fewer creators can afford to produce high-quality content. The industry is watching closely to see how this legal battle and the internal reflections within AI giants will shape the future of information access and AI development.