Pending AI training cases: a user's risk map
NYT v. OpenAI, the Disney and Warner suit against Midjourney, and the Anthropic lyrics cases are all in discovery. What a downstream user should infer.
The three US cases most often cited as proof that AI content is legally radioactive have one thing in common: none of them has produced a merits ruling. They are in discovery. Documents are being exchanged. Experts have not reported. No judge has said who is right.
That is not a reason to relax. It is a reason to be precise about where your actual exposure sits, because the answer is not where the headlines put it.
The map, as of August 2026
| Case | Court | Posture | Next known milestone | What it is really about |
|---|---|---|---|---|
| NYT v. OpenAI and Microsoft | S.D.N.Y. | Discovery; a sanctions motion was filed Jul 2026 | No trial date | Training on news archives, plus alleged regurgitation of article text |
| Disney, Universal and Warner Bros. v. Midjourney | C.D. Cal., consolidated 4 Nov 2025 | Discovery | Expert disclosures Oct 2026 | Outputs depicting protected characters |
| Concord / UMPG / ABKCO v. Anthropic | — | Two live suits, no settlement | — | Reproduction of song lyrics |
| UMG v. Suno | — | Continuing; Suno defending on fair use | — | Music training. Sony has settled with neither Suno nor Udio |
Everything in that table is pending. Read the "posture" column twice before you read anything else.
What "in discovery" actually means
A case in discovery has cleared the pleading stage and is now in document exchange. After that come expert reports, then summary judgment briefing, then possibly trial, then very possibly appeal. Two to four years is unremarkable.
Two things get misreported as substantive developments:
- A sanctions motion is a discovery-conduct dispute. It is an argument about how a party has behaved in the process. It has no bearing on whether training was fair use.
- An expert disclosure deadline is a calendar entry. It tells you the case is moving on schedule. It does not tell you what the experts will say or whether the court will credit them.
The general rule: a pending case generates coverage proportional to the size of the plaintiff and information proportional to zero. When someone tells you the law has changed, ask which court, which date, and whether the decision is on the merits.
The one case shaped like a user's problem
Look again at the Midjourney row. The claim is about outputs depicting protected characters.
That is a structurally different animal from a training claim, and the difference is the whole point of this post:
- Training claims are about what the provider did before you arrived. You are not a party, you have no control, and a loss by the provider does not create a claim against you.
- Output claims are about what came out. There is a natural downstream defendant in an output claim, and it is the person who typed the prompt, exported the file, and put it in a paid ad.
Nobody has ruled on the studios' theory yet. But it is the theory with a path to your desk, and it is the one worth building habits around.
Two contract facts that people get backwards here:
- A provider can only assign what it holds. Terms that grant you the provider's right, title and interest in the output do not create copyright where the human-authorship test fails, and they do not extinguish anyone else's rights in the material the output resembles.
- Terms of use do not immunise you against third-party claims. Read your plan's indemnity clause specifically. Copyright indemnities from the major providers are generally enterprise or API tier, are usually conditioned on leaving the safety filters on, and generally are not offered on consumer plans at all. Assume you have none unless you can quote the clause.
What to infer, and what not to
Do not infer:
- That a pending case means the practice is unlawful. It means somebody alleged it is.
- That a settlement elsewhere resolves anything. The Anthropic authors' settlement created no precedent, and the music settlements — Universal with Udio, Warner with Udio and with Suno — were commercial deals paired with forward licences, not merits rulings.
- That "no ruling yet" equals safe. Discovery is where the expensive facts surface.
- That a sued provider is unusable. Every large provider in this category is a defendant somewhere.
Do infer:
- That reproducing recognisable third-party IP in a commercial deliverable is the live, unresolved risk, and it is the one a plaintiff can trace to you.
- That music has settlements and forward licences but still no merits ruling on training. A signed label deal is not a cleared model is the version of this that matters if you are scoring ads.
- That your indemnity tier is a real variable and most people have never checked theirs.
A posture that survives any of these outcomes
- Do not prompt named characters, franchises, or trade dress for commercial work. This is not caution, it is the exact claim currently in discovery in California. Style references are a different question; a recognisable character is not a style.
- Run an output check before delivery, not after. Logos, marks, faces, lyrics, and anything that looks borrowed. The brand safety checklist is the version to hand a junior.
- Know why a model refused you. Refusal categories are a decent proxy for where providers think their own exposure is. Reading a model's safety policy before you write the prompt saves both the generation and the argument.
- Keep a generation record per deliverable. Model, date, prompt, references supplied, whose likeness appears, what was licensed. This is the single cheapest insurance in the whole workflow and almost nobody does it.
- Write the rights into the contract. Scope, territory, term, and who carries what if a claim lands. The practical shape is in legal and licensing for AI content in business.
- Watch the distribution layer, not just the courts. Marketplace and platform rules move in weeks rather than years, and they will block your listing long before a judge does. Where generated product shots are banned is the current state of that.
One more selection input, upstream of all of it: what a provider publishes about its own data. That is knowable today and does not depend on any of these dockets. Training-data provenance as a buying criterion covers how to weigh it against everything else you are choosing on.
FAQ
Is an individual creator likely to be sued over a generation?
The active suits target model providers, not users, and the economics of suing an individual are poor. The realistic downstream consequences are commercial rather than judicial: a client refusing delivery, a platform removing a listing, an ad network rejecting creative, or a brand's legal team killing a campaign. Those happen regularly and they happen fast.
If none of these cases has ruled, is the current practice fine?
No. Absence of a ruling is absence of information, not permission. It also cuts the other way — a plaintiff who has not won has also not lost, and the studios' output theory is being built right now with the benefit of discovery.
If a provider settles, does that clear its model for my use?
No. A settlement binds the parties to it. It may come with a forward licence that improves the provider's position going forward, which is what happened in the music settlements, but it does not adjudicate anything and it does not transfer a clearance to you. For music work, licensing and royalty questions are still decided by what you can document, not by who settled.
Does a licensing deal between a studio and an AI company change my position?
Only if you are inside its scope, which you almost certainly are not. Those deals are between two companies and can be withdrawn or collapse; what Disney's collapsed OpenAI deal signals is a useful case study in how quickly a headline arrangement can stop existing. Do not treat a press release as a clearance.