Anthropic · OpenAI · Donald Trump · New York · ChatGPT · TechCrunch AI
The LLMs powering chatbots like ChatGPT, Claude, and Gemini are trained on incomprehensibly massive databases of published works
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
This question, can you use copyrighted material to train an AI?, isn’t black and white, hence the extensive legal debate around the subject.
Key facts
- Last year, Judge William Alsup ordered Anthropic to pay a $1.5 billion copyright settlement to a group of writers whose works were used to train the company’s AI models; but Anthropic wasn’t dinged
- In a lawsuit that The New York Times filed against OpenAI, the Trump administration has contributed a 20-page brief in defense of the ChatGPT maker’s unlicensed use of copyrighted material to train
- This new Trump administration brief is not a ruling, as the case is being tried in the U.S. District Court for the Southern District of New York, and the authors of the brief do not have jurisdiction
- So far, cases about AI training and copyright infringement have largely been favorable to AI companies
Summary
In a lawsuit that The New York Times filed against OpenAI, the Trump administration has contributed a 20-page brief in defense of the ChatGPT maker’s unlicensed use of copyrighted material to train its LLMs. “The United States has a strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally… As such, it is critical for the United States to ‘retain global leadership in artificial intelligence,’” the brief reads, referencing an executive order that President Donald Trump signed last year. The LLMs powering chatbots like ChatGPT, Claude, and Gemini are trained on incomprehensibly massive databases of published works, including copyrighted books, articles, and other media that AI companies feed into these databases without permission.