The training data argument has produced three years of essays and almost no money changing hands.
That ended. Anthropic's copyright settlement puts $1.5 billion into a pool that pays roughly $3,000 per pirated work, and the claims process is now running.
The interesting conflict is not the one anybody was expecting.
The fight moved inside the industry
TechCrunch reported this week that authors are pushing back as publishers and literary agents make claims on the settlement money.
The split is defined. For in-print books, $3,000 is divided fifty-fifty between author and publisher. For self-published works, or works whose rights reverted to the author, the author takes 100 percent.
Reversion is the pressure point, and it turns on a date. Rights must have reverted before 10 August 2022, the settlement's download date, for an author to claim the full amount.
According to the reporting, publishers including HarperCollins have filed claims on books whose rights reverted years ago, in some cases seeking the full payment where they would be entitled to half.
Literary agencies are filing claims too, which is harder to explain. As author Courtney Milan put it, agents are not rights holders in the books that they sell.
Victoria Strauss, who tracks this territory, reports the same errors appearing repeatedly rather than as isolated mistakes, which points at something systemic rather than clerical.
Objection only pays if you can document ownership
Strip away the specifics and there is a structural lesson here for anyone whose content is an input to a model.
Authors are the group that successfully objected. Not by writing about it, by litigating, and the objection converted into $1.5 billion.
Then the money arrived and it turned out that objecting was the easy half. The hard half is proving you are the party entitled to it, using paperwork drafted decades before training data was a concept anybody had.
Most publishing contracts do not contain language about machine training. They contain language about print, digital, audio, and territories. So the question of who owns the training-use payout gets resolved by analogy, and analogies get resolved in favour of whoever has counsel on retainer.
The same gap sits inside almost every content contract signed in the last twenty years, including yours.
Ghostwritten thought leadership. Agency-produced content where the agency retained rights. Photography licensed for web use in 2014. Guest posts. Employee-authored material where the employment contract predates any of this.
In each case somebody will eventually own the training-use claim, and it will be determined by contract language written for a different purpose.
Meanwhile the lawsuits keep arriving
The same week, The Seattle Times and Newsday filed suit against OpenAI and Microsoft in the Southern District of New York, alleging unauthorised use of their journalism for training.
The filing describes the AI products as consumers of human-authored content, producing derivative copies for commercial gain. It follows a pattern that started with The New York Times in December 2023 and has not slowed.
One detail is worth noting for how tangled this has become. Microsoft and OpenAI had previously funded journalism projects at The Seattle Times.
Funded by them, now suing them. That is not hypocrisy, it is what happens when a licensing market has no price and no standard terms, so money moves as sponsorship in one direction and damages in the other.
I wrote in an earlier edition about the AI content bill coming due. This is that bill being itemised, and the itemisation is where the real disputes live.
$3,000 becomes the anchor
A settlement figure does something a court opinion does not. It creates a reference price.
Until this week, the question of what a book is worth as training data had no public answer, so every licensing negotiation started from nothing and ended wherever the leverage sat. Now there is a number, and numbers are sticky.
Look at what it implies. A $1.5 billion pool at roughly $3,000 per work prices the pool at around half a million works.
Set that against what models trained on those works have earned and the figure is not compensation for value created. It is the cost of extinguishing a liability, which is a different calculation and a much smaller one.
Everyone negotiating content licensing over the next two years will have this number quoted at them, by whichever side it helps. Rights holders will call it a floor. Buyers will call it a benchmark.
The harder problem is that a per-work price does not translate to most content. What is a product page worth? A case study? A five-year archive of blog posts?
There will be no equivalent class action for marketing content, because the damages per item are too small for anyone to organise around. Which means for most companies, the training-use question never gets settled in a courtroom at all.
It gets settled by your crawler directives, your contracts, and whatever you are able to enforce on your own.
What to do with your own archive
Two things, and the first is cheap.
Find out who owns the content you have published. Not who wrote it, who holds the rights. For agency-produced work, check whether the contract assigned copyright or licensed it, because the difference decides whether any future claim is yours or theirs.
For freelance and contractor work, check whether you have work-for-hire language or a licence, and check whether the licence enumerates uses. An enumerated licence probably does not cover training, which means the contractor may hold that right.
Second, put the language in new contracts now. One clause specifying who holds rights to use of the work as training data for machine learning systems, and who receives proceeds from any collective claim or settlement.
It costs nothing to include today. In four years it will be the clause everybody wishes they had written, in the same way that everybody currently wishes their 2015 contracts had said something about digital archives.
The authors did the hard part and won the argument. What is happening now is the part nobody prepares for, which is that winning creates an asset and assets attract claimants.
Your archive is going to be worth something. Find out whose.