A ghostwriter uses your grandmother's diary to draft a family memoir, sells it, and never tells you. You would not need a lawyer to know something went wrong. That instinct, felt by millions of writers when they realized their books were inside GPT and Claude, is the starting point for a legal fight that has barely begun. At YouWrite we build on top of models trained this way, so the honest thing is to explain what happened rather than dress it up.
What 'training data' actually is
A large language model is not a database of books. It is a set of statistical weights, billions of numbers, tuned by exposure to text. During training, the model reads a sequence and adjusts itself to predict the next word better. Do this trillions of times across a huge corpus and you get something that can write in the register of Zadie Smith or the cadence of a court filing.
The corpus matters. For the earliest GPT models, OpenAI used a filtered slice of Common Crawl, a nonprofit web scrape, plus WebText (Reddit-linked pages) and two book collections labeled Books1 and Books2. Books2 has never been publicly identified by OpenAI. In the Authors Guild v. OpenAI complaint filed September 2023, the plaintiffs alleged Books2 was sourced from a shadow library like LibGen or Z-Library. In 2024, The Atlantic published a searchable index showing that Meta's Llama models were trained on LibGen. Meta acknowledged this in filings for Kadrey v. Meta.
So when writers say 'they took my book,' this is often literally true at the ingestion step. What happens after is contested.
Scraping, in plain terms
Scraping means a program requests web pages the way a browser does, then saves the text. Common Crawl has done this openly since 2008 and publishes the archive. Anyone can download it. The scrape itself is not the controversial part. The controversy sits in what gets included (pirated book dumps, paywalled journalism, fanfiction archives that scraped their own users) and what the resulting model does with it.
Why copyright law is genuinely stuck
The reflexive answer is 'this is theft.' The reflexive counter is 'this is fair use, like a human reading a library.' Both are too easy.
US copyright law has a four-factor fair use test. The one everyone argues about is factor one, whether the use is 'transformative.' In Authors Guild v. Google (2015), the Second Circuit held that Google's scanning of millions of books to build a searchable snippet index was transformative fair use. That decision is the closest existing precedent, and OpenAI cites it constantly. Some serious scholars, including Pamela Samuelson at Berkeley, have written that AI training has a plausible transformative-use argument, even while criticizing how the industry has behaved. Matthew Sag at Emory has argued for years that machine learning on text is generally lawful under a doctrine he calls 'non-expressive use.'
The counterarguments are also serious. A search index returns snippets and points you back to the book. A generative model can produce prose in an author's voice, sometimes reproducing training text verbatim, and competes in the same market as the author. The New York Times v. OpenAI complaint (December 2023) included dozens of examples of GPT-4 reproducing Times articles nearly verbatim from prompts. That is not a snippet. That is the fourth fair use factor, market effect, staring the court in the face.
The honest position: fair use here is unsettled. Anyone telling you it is obviously legal, or obviously not, is selling something.
What is realistically going to happen
Three tracks are already visible.
Lawsuits will produce narrow rulings. The Times case, Kadrey v. Meta, Silverman v. OpenAI, Getty v. Stability, and the Concord Music suits against Anthropic will not deliver one grand verdict. They will produce fact-specific holdings about memorization, market harm, and whether pirated sources poison the fair use analysis. In February 2025, Judge Bibas in Thomson Reuters v. Ross Intelligence ruled against Ross on fair use, the first substantive US ruling on AI training and copyright. It involved a legal research tool, not a generative model, but it signaled that courts will look hard at market substitution.
Licensing deals will normalize. OpenAI has signed with the Associated Press, Axel Springer, News Corp, the Financial Times, Vox Media, and Reddit. Google has a reported $60 million deal with Reddit. These deals establish a market price, which quietly undermines the fair use argument for holdouts. If AP is worth licensing, why is your novel not?
Opt-out registries will spread and mostly disappoint. The Have I Been Trained tool from Spawning lets you check if your work is in LAION. OpenAI's Media Manager was announced in May 2024 and has not shipped as of this writing. Opt-out puts the burden on the writer and only affects future training runs, not the models already trained on you.
What writers can actually do this week
- Register your work with the US Copyright Office if you are American, or your national equivalent. Statutory damages are only available for registered works, and you cannot join most class actions without registration. This is dull and it matters.
- Join a body that can bargain collectively. The Authors Guild has been the lead plaintiff in the OpenAI suit. The National Union of Journalists in the UK has been active on training data. The Society of Authors has pushed for licensing frameworks. Individual protest letters do not move a $150 billion company. A union that can coordinate 10,000 opt-outs and a lawsuit does.
- Add a robots.txt disallow for GPTBot, ClaudeBot, CCBot, and Google-Extended if you run your own site. It will not undo past training. It signals non-consent for future scrapes, which matters legally and practically.
- Watch the Spawning registry and OpenAI's Media Manager when it launches. Register there too. Redundancy is fine.
The real issue is consent, and the law has not named it yet
Writers are not confused when they feel wronged. A commercial product was built from their labor without asking. Copyright is the tool at hand, but it was designed for a world of printing presses and photocopiers, not statistical models that learn style. The right framework may end up looking more like the moral rights tradition in French law, or the collective licensing that ASCAP built for songwriters after radio broke the old model.
We use these models at YouWrite to help writers draft long-form story work, and we tell clients openly what the models are and where they came from. If you want to know whether your book is in the training set, check Spawning and the Atlantic's LibGen index. If you want to change what happens next, register your copyright, join the Authors Guild, and block the crawlers by Friday.
