Copilot · OpenAI · Microsoft · New York · The Verge
Microsoft confirms virtually nobody was grabbing NYT articles through its chatbot
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
Fewer than 1 percent of more than 8 million chat logs regurgitated at least 16 words.
Key facts
- Fewer than 1 percent of more than 8 million chat logs regurgitated at least 16 words
- Similarly, an expert in the authors’ suit found that the 8.2 million conversations with Copilot only had 24 responses that contained at least 30 matching words
- An expert for the Center for Investigative Reporting found 51 instances of “substantial overlap” with CIR work in the dataset, Microsoft says
- As part of the lawsuit’s discovery, Microsoft provided 8.2 million Copilot chat logs to an expert hired by news publishers
Summary
Microsoft’s Copilot rarely reproduces even full sentences from news articles and books, let alone substantive chunks that could substitute for the original, the company says in new legal filings as it fights copyright claims from publishers including The New York Times and book authors. As part of the lawsuit’s discovery, Microsoft provided 8.2 million Copilot chat logs to an expert hired by news publishers. Similarly, an expert in the authors’ suit found that the 8.2 million conversations with Copilot only had 24 responses that contained at least 30 matching words. The Times disagreed with Microsoft’s conclusions. Microsoft argues that the numbers bolster its case that using copyrighted content for AI training datasets should be considered fair use.