A wave of copyright claims from authors and newspapers is forcing the biggest technology companies to defend how their models were trained, with Anthropic reportedly agreeing to a landmark settlement and the New York Times case against OpenAI still unresolved.
For years the artificial intelligence boom has been powered by a single, unglamorous ingredient: enormous quantities of text and images scraped from across the internet and beyond. Now, in courtrooms across the United States, the question of exactly where all that material came from has quietly become one of the industry's most expensive problems.
A growing wave of copyright lawsuits brought by authors, newspapers and artists is forcing the world's leading technology companies to defend how their models were actually built. The outcomes, still unfolding through 2026, could reshape the economics of an entire field and settle who really owns the raw material behind these systems.
The Landmark Settlement
The most striking development so far came from Anthropic, the company behind the Claude chatbot. According to reports, the firm agreed to pay authors roughly 1.5 billion dollars to resolve a class-action lawsuit, in what has been widely described as one of the largest copyright settlements in American history.
The sheer scale of that figure sent an unmistakable signal across the sector. For an industry that had often treated training data as a free and effectively limitless resource, a payout of that size to a group of writers served as a sobering reminder that the words feeding these models belong to somebody in the first place.
How the Case Reached a Settlement

The dispute did not turn simply on whether software can learn from books at all. Earlier in the case, Judge William Alsup reportedly found that training a model on lawfully acquired text could count as a transformative fair use, a conclusion that was seen at the time as broadly encouraging for the technology companies.
The real problem lay elsewhere. The same ruling reportedly allowed the case to proceed toward trial over how some of the books had been obtained, with the plaintiffs arguing that many titles had been downloaded from pirate libraries rather than bought. That distinction, between learning from a work and taking it unlawfully, proved decisive.
The Terms of the Deal
The reported settlement is said to cover a vast catalogue of writing. According to reports, it applies to somewhere around 500,000 books, with the company agreeing to pay in the region of 3,000 dollars for each covered title, a sum that could shift depending on how the final accounting is eventually handled by the court.
The agreement reportedly comes with conditions that reach beyond the money itself. Anthropic is said to have agreed to destroy the pirated files it had accumulated, a term that underlines how the case was less about the act of training a model and far more about the tainted origins of the underlying data.
The Times Takes On OpenAI
If the Anthropic case offered a glimpse of one possible ending, the industry's most closely watched legal battle remains firmly unresolved. The New York Times sued OpenAI and its partner Microsoft back in December 2023, accusing them of using millions of its articles without permission to help build their popular systems.
More than two years on, that case is still grinding steadily forward. It has moved deep into the discovery phase before a federal court in New York, where lawyers have spent long months examining how the models were assembled and precisely what material was fed into them along the way to completion.
A Ruling With Consequences
Both sides have now reportedly asked the court to rule in their favour before any trial, filing competing motions for summary judgment. A decision is expected later this year, and few observers seriously doubt that it will echo far beyond the two companies directly involved in this particular fight.
For news publishers in particular, the stakes are enormous. A ruling that favours the newspaper could hand the wider media industry powerful leverage to demand payment, while a decision for the technology company could confirm that training on published journalism sits comfortably within the existing boundaries of the law.
What Comes Next
For now, no settlement between the newspaper and OpenAI has been reported, and the two remain locked in a contest that could define the limits of fair use for a generation. Other authors, musicians and film studios are watching the proceedings closely, quietly weighing their own potential claims against the giants of the field.
What began as an abstract debate about how machines learn has hardened into a concrete fight over money, consent and credit for creative work. However the remaining cases are ultimately resolved, the era in which training data could simply be taken without real consequence now appears, slowly, to be drawing to a close.
Balanced view on copyright.
Nice deep look at copyright.

Keep following Ethan BrooksHer next filing reaches you the moment it publishes, on her own subdomain.
Follow