The missing ingredient
A ramen shop in Kamakura, the lawsuits no one wanted to pay for, and why AI's data problem is a procurement problem the law was never meant to solve.
Treating training data as a free input was never free. The bill was deferred, not waived. Every settlement we have read about in the last three years is the same invoice arriving late, in a courtroom, at a multiple of what licensing would have cost up front.
There is a small ramen shop in Kamakura called Awanouta. The pork comes from a farm an hour up the coast. The vegetables come from a market two stops down the train line. The broth is the shop’s own work, but every ingredient that enters the kitchen has been bought, by someone, from someone, on the way in.
A bowl of ramen, if you look at it the right way, is a procurement document. It is also, at the moment, a better procurement document than most AI labs can produce.
The analogy is the whole point
The current generation of frontier models was trained on something. That something includes:
- songs scraped from SoundCloud and YouTube
- artwork scraped from DeviantArt, ArtStation, and Pinterest
- code scraped from GitHub, including private gists that leaked
- pirated PDFs from shadow libraries (Books3, LibGen, Z-Library)
- millions of blogs, essays, comments, transcripts, and captions
- the work of people who are alive, contactable, and uncompensated
Some of this was scraped under unsettled doctrine. Some of it was scraped knowing the doctrine was not unsettled at all. None of it was procured the way ingredients are procured.
A small comparison:
| Ramen shop | AI lab |
|---|---|
| Buys ingredients from farmers | Should buy data from creators |
| Pays a fair market rate | Pays nothing, or settles later |
| Builds long-term supplier relationships | Treats suppliers as adversaries |
| Cooks meals with paid inputs | Trains models on unpaid inputs |
| Customers pay for the experience | Customers pay for the output |
The asymmetry is not a quirk. It is the entire business model of the early AI era.
The bill is already arriving
When I first wrote about this in 2025 I had one big case to point to: Sarah Silverman and a handful of authors suing OpenAI and Meta in 2023 for training their LLMs on pirated book corpora pulled from shadow libraries. The original reporting was in The Verge, and the underlying scraping of Books3 (a 196,640-book dataset assembled from Bibliotik torrents) was documented by The Atlantic the same year.
Since then the docket has gotten heavier and the rulings have gotten more concrete.
The New York Times v. OpenAI and Microsoft, filed in December 2023, has moved from filing to substantive arguments about reproduction, licensing harm, and the discoverability of training inputs. Bartz v. Anthropic produced a settlement large enough to be itemized in the lab’s reporting. Disney and NBCUniversal filed against Midjourney in June 2025 with image grids showing near-exact reconstructions of their characters. Universal Music Group and the other majors moved on Suno and Udio in mid-2024 and have been settling and re-filing in waves since.
The pattern is consistent.
Labs train. Discovery happens. The corpus contains things it should not contain. The cases settle, or the cases go to verdict, and the cost of having skipped the procurement step lands back on the lab’s balance sheet at a multiple of what licensing would have cost up front.
This is a procurement problem, not a copyright problem
The legal frame is doing the wrong job. Copyright is a backstop, the rule of last resort when no contract exists. The reason copyright is doing so much work in AI right now is that we never built the contract layer underneath it. There is no clean API that lets a model say “I would like to use this work, in this context, for this fee,” and get an answer in time to matter.
Until that layer exists, every training run is a procurement decision made in the absence of a procurement system. Some teams make it cynically. Some make it carelessly. The structural answer is not better cynicism, and not louder carelessness. The structural answer is a market.
The ramen shop in Kamakura is not a metaphor for the future of AI. It is a working example of a kind of supply chain we used to take for granted. Inputs were paid for. Suppliers were respected. The price of the meal reflected the price of the produce.
What we are building
IPTO is the bet that the licensing layer is buildable, that creators will price their work into a clear market once one exists, and that labs will prefer to pay a known number up front than discover the unknown number in a courtroom three years later. The rails for that are what we are shipping.
Human Data Rights does the upstream work: making sure the rights themselves are legible, transferable, and revocable in a way the market can use. A licensing layer is only as honest as the property layer underneath it.
The shorter version: AI is here to stay. The only open question is whether we feed it the way Awanouta cooks, or the way a thief does.
Sources
- Silverman et al. v. OpenAI, Inc. and Silverman et al. v. Meta Platforms, Inc. (N.D. Cal., filed July 2023). Coverage: The Verge.
- The New York Times Company v. Microsoft Corp. & OpenAI (S.D.N.Y., filed December 2023). The Times’s own coverage and complaint.
- Bartz et al. v. Anthropic PBC (N.D. Cal., 2024). Settlement reported across major outlets in 2025.
- Disney Enterprises & NBCUniversal v. Midjourney, Inc. (C.D. Cal., filed June 2025).
- UMG Recordings et al. v. Suno, Inc. and v. Uncharted Labs (Udio) (D. Mass. and S.D.N.Y., filed June 2024).
- The Books3 dataset and its provenance in pirated Bibliotik torrents, documented by The Atlantic and Wired in 2023.
- The ramen, every time, from Awanouta in Kamakura.