Microsoft Argues Copilot Chatbot Rarely Copies Copyrighted Content
Defending itself in copyright lawsuits filed by publishers and authors, Microsoft stated that its Copilot artificial intelligence tool almost never copies original news and book content.
In new legal filings regarding copyright lawsuits brought by The New York Times and other publishers, Microsoft stated that the Copilot chatbot rarely reproduces news and book content in full sentences.
Legal Filing Details in the Proceedings
As part of its fight against copyright claims, Microsoft submitted new legal filings arguing that Copilot does not copy significant portions that could serve as substitutes for original works.
Chat Logs and Analysis Results
During the discovery process, Microsoft provided 8.2 million Copilot chat logs to an expert retained by publishers. It was stated that these logs were selected using keywords pointing to the usage of publishers' websites.
As a result of the analyses, it was determined that a small fraction of the records had an overlap of at least 16 words with the news content used to support the AI model.
Findings by Publishers and Experts
An expert working for the Center for Investigative Reporting stated that they found 51 examples in the dataset that significantly overlapped with the organization's work. Another expert in the authors' lawsuit indicated that only 24 out of 8.2 million conversations contained at least 30 matching words.
The New York Times's Response
The New York Times objected to these results, while the lead attorney for the lawsuit, Ian Crosby, argued that the emerging documents show Microsoft and OpenAI took journalistic products to build commercial products.
Fair Use Defense and Expectations
Microsoft argues that these numbers reinforce its arguments that the use of copyrighted content in AI training datasets should be considered fair use. The filing was made with a request for summary judgment to dismiss the case at an early stage.