Don't Steal This Book: 10,000 Authors Publish Empty Novel to Protest AI Copyright Theft
Independently fact-checked against primary sources (last audited August 24, 2026). · Reviewed by the RecordingLaw editorial team. · Law checked current as of August 24, 2026. · 5 primary sources cited on this page. How we verify our legal content

Nearly 10,000 authors, including Nobel laureate Kazuo Ishiguro, published a book with 88 pages of names followed by nothing but blank pages. Distributed at the London Book Fair in March 2026, "Don't Steal This Book" is a protest against AI companies that scraped copyrighted creative works without permission or payment to train their models.
The stunt landed during a critical window. The UK government was weighing whether to upend copyright law in favor of AI companies, a music piracy lawsuit seeking more than $3 billion had been filed against Anthropic, and Anthropic itself had accused three Chinese AI labs of stealing its own model outputs. The message from the creative community was blunt: the AI industry is built on stolen work.
The Empty Book
"Don't Steal This Book" is a physical object designed to make an argument. The first 88 pages list the names of the contributing authors. After that, every page is blank.
The symbolism is intentional. If AI companies continue to train large language models on copyrighted novels, short stories, and other creative writing without licensing or compensation, the blank pages represent the future of authorship: writers who can no longer sustain careers because machines trained on their work have replaced the market for it.
The book was organized by Ed Newton-Rex, a composer and technologist who left his role at an AI company after growing disillusioned with how the industry treated creative intellectual property. Newton-Rex founded Fairly Trained, an organization that certifies AI models built on properly licensed data.
"The AI industry is built on stolen work, taken without permission or payment," Newton-Rex told The Guardian. He described the book as a direct message to the UK government, which was then considering changes to copyright law that could benefit AI developers at the expense of creators.
Among the nearly 10,000 contributing authors were Kazuo Ishiguro (Nobel Prize in Literature, 2017), Richard Osman, Alan Moore, Marian Keyes, Philippa Gregory, Malorie Blackman, and Mick Herron. Approximately 1,000 physical copies were distributed across the London Book Fair, held March 10-12, 2026.
The UK Copyright Battle
The protest's timing was strategic. The UK government was required to submit both an economic impact assessment and a comprehensive progress report on copyright and AI to Parliament by March 18, 2026, under Section 137 of the Data (Use and Access) Act.
What the Government Proposed
In December 2025, the UK government had identified a broad data mining exception with an opt-out mechanism as its preferred approach. Under this framework, AI companies could train models on any lawfully accessed copyrighted material, and rights holders who objected would need to actively opt out.
How Creators Responded
The proposal was overwhelmingly rejected by the creative industries. During the public consultation (December 2024 through February 2025), 88% of over 11,500 respondents supported requiring licenses in all cases. Only 3% backed the government's preferred opt-out model.
A taskforce of 80 people manually analyzed all responses without using AI tools. Over 3,000 of the responses were based on template letters organized by creative industry groups, reflecting coordinated opposition.
The Government's Reversal
On March 18, 2026, the UK government announced it had reversed course. It withdrew support for the opt-out framework and stated it now has "no preferred option" for reform. The government committed to continuing evaluation of policy options, monitoring international developments, and exploring licensing mechanisms for smaller organizations.
Three core principles were outlined for any future framework: rights holders should control how their work is used and receive compensation, developers need lawful access to content for training, and transparency requirements must give visibility into what data models are trained on.
The Anthropic Music Lawsuit
While authors protested in London, the courtroom fight over AI and copyright was escalating in the United States.
In January 2026, a coalition of music publishers led by Universal Music Group and Concord Music Group filed suit against Anthropic seeking more than $3 billion in damages. The complaint alleges Anthropic illegally downloaded over 20,000 copyrighted songs, including sheet music, lyrics, and musical compositions, to train its Claude AI model.
According to the filing, Anthropic built a permanent internal library of copyrighted texts sourced from so-called pirate repositories rather than acquiring licensed copies. The case evolved from an earlier, smaller lawsuit covering approximately 500 works. Through the discovery process, the publishers claim they found evidence of a far larger operation.
If the damages are awarded in full, the case would rank as one of the largest non-class action copyright cases in U.S. history.
The Bartz Settlement and Fair Use Precedent
The music lawsuit follows the landmark Bartz v. Anthropic settlement, which resolved a class action brought by authors. In that case, Judge William Alsup issued a critical distinction that has shaped the legal landscape: training AI models on copyrighted content may qualify as fair use, but acquiring that content through piracy does not.
The case settled for $1.5 billion, with impacted writers receiving approximately $3,000 per work for roughly 500,000 copyrighted works. The ruling established an important boundary: AI companies can argue fair use for the training process itself, but they must obtain their training data through legitimate channels.
This distinction matters because it shifts the legal battleground. The question is no longer just "can AI companies use copyrighted works to train models?" but "how did they acquire those works in the first place?"
When AI Companies Get Stolen From
In a twist that underscored the complexity of intellectual property in the AI era, Anthropic itself accused three Chinese AI laboratories of stealing its work.
In February 2026, Anthropic publicly alleged that DeepSeek, Moonshot AI, and MiniMax had orchestrated industrial-scale "distillation" campaigns targeting Claude. The three labs allegedly created approximately 24,000 fraudulent accounts and generated over 16 million exchanges designed to systematically extract Claude's reasoning capabilities.
MiniMax was the most prolific, generating over 13 million exchanges focused on coding and tool use. Moonshot AI accounted for more than 3.4 million exchanges. DeepSeek's operation was smaller (over 150,000 exchanges) but strategically targeted Claude's reasoning and reinforcement learning capabilities.
The irony was not lost on commentators: an AI company accused of building its models on pirated creative works was itself alleging that foreign competitors had pirated its AI outputs. No DOJ investigation into the distillation allegations has been confirmed. The corroborated federal response came in the following weeks: on April 15, 2026, Representatives Bill Huizenga and John Moolenaar introduced the Deterring American AI Model Theft Act of 2026 (H.R. 8283), which would direct Commerce Department Entity List designations and authorize sanctions under the International Emergency Economic Powers Act against foreign entities that misappropriate US AI models if enacted, and on April 23, 2026, the White House Office of Science and Technology Policy issued a national security memo accusing China of running industrial-scale AI distillation campaigns and directing federal agencies to share intelligence with AI companies. As of August 2026, the bill remains pending in Congress.
The Legal Landscape for AI and Copyright
The fight over AI training and copyright is playing out across multiple legal fronts simultaneously.
U.S. Fair Use Uncertainty
Courts have split on whether AI training constitutes fair use under 17 U.S.C. 107. The Bartz ruling allowed fair use for training but drew the line at pirated source material. In Kadrey v. Meta Platforms, the court noted that AI models can "flood the market with similar texts and stifle competition," identifying potential market harm that could undermine fair use claims.
On March 2, 2026, the U.S. Supreme Court denied certiorari in Thaler v. Perlmutter, confirming that AI-generated content cannot be copyrighted without human authorship. This ruling establishes that while AI companies may (or may not) have the right to use copyrighted works as training data, the outputs those models generate do not themselves receive copyright protection.
International Approaches
Countries are taking divergent paths:
| Jurisdiction | Approach |
|---|---|
| United States | Case-by-case fair use analysis; no comprehensive AI copyright legislation |
| United Kingdom | Reversed preferred opt-out model; now has "no preferred option" |
| European Union | Text and data mining exception under DSM Directive, with opt-out for commercial use |
| Japan | Broad exception for AI training under 2018 copyright amendments |
The lack of international consensus means AI companies operate under different rules depending on where they and their training data are located. This creates both legal uncertainty and opportunities for jurisdiction shopping.
What This Means for Digital Rights
The "Don't Steal This Book" protest and the lawsuits surrounding it sit at the intersection of copyright, data privacy, and digital rights. Several principles are emerging from the legal battles.
Consent matters. Whether it is authors objecting to their books being scraped, musicians objecting to their songs being pirated, or users objecting to their recordings being reviewed by contractors, the common thread is that people and organizations want control over how their creative output is used.
Piracy is still piracy. The Bartz settlement drew a bright line: however courts ultimately resolve the fair use question for AI training, obtaining training data through pirated or unauthorized channels is not protected. This principle applies equally to the Anthropic music case and to the DeepSeek distillation allegations.
Transparency is the minimum. Both the UK government's revised framework and the EU's DSM Directive emphasize transparency requirements. AI companies may eventually be required to disclose what copyrighted works they used for training, giving rights holders the information they need to enforce their rights.
The blank pages of "Don't Steal This Book" ask a question that legislatures, courts, and AI companies have not yet fully answered: who gets to profit from creative work, and who decides?
Consult an attorney for advice specific to your situation, particularly if your copyrighted work may have been used in AI training without authorization.
Frequently Asked Questions
What is 'Don't Steal This Book'?
It is a mostly empty book published by nearly 10,000 authors, including Kazuo Ishiguro, Richard Osman, and Alan Moore. The first 88 pages list the contributing authors' names, followed by blank pages symbolizing the future of authorship if AI companies continue to train on creative works without permission. 1,000 copies were distributed at the London Book Fair in March 2026.
Why are authors protesting AI companies?
Authors allege that AI companies scraped copyrighted novels, stories, and other written works to train large language models without obtaining licenses or paying royalties. Organizer Ed Newton-Rex described the AI industry as 'built on stolen work, taken without permission or payment.' The protest targeted the UK government's consideration of copyright changes that could benefit AI developers.
What is the Anthropic music lawsuit about?
In January 2026, music publishers led by Universal Music Group and Concord Music Group sued Anthropic for over $3 billion, alleging the company illegally downloaded more than 20,000 copyrighted songs to train Claude. The publishers claim Anthropic built an internal library from pirate repositories rather than acquiring licensed copies.
Is it legal for AI companies to train on copyrighted works?
The answer depends on the jurisdiction and how the works were acquired. In the U.S., the Bartz v. Anthropic case established that AI training may qualify as fair use under 17 U.S.C. 107, but obtaining training data through piracy is not protected. Courts remain split on the broader fair use question. The UK has no settled framework yet.
What did the UK government decide about AI copyright?
On March 18, 2026, the UK government reversed its preferred opt-out approach (where AI companies could train on any lawfully accessed material unless rights holders opted out). After 88% of over 11,500 consultation respondents opposed this, the government stated it now has 'no preferred option' and will continue evaluating alternatives.
What happened with DeepSeek and Anthropic?
In February 2026, Anthropic accused three Chinese AI labs (DeepSeek, Moonshot AI, and MiniMax) of creating approximately 24,000 fraudulent accounts to extract Claude's capabilities through over 16 million exchanges. No DOJ investigation has been confirmed; the corroborated response came weeks later, when the White House issued a national security memo (April 2026) accusing China of industrial-scale AI distillation and Congress introduced the Deterring American AI Model Theft Act of 2026 (H.R. 8283), which would direct Entity List designations and IEEPA sanctions against foreign AI-model theft if enacted.
Can AI-generated content be copyrighted?
No, under current U.S. law. On March 2, 2026, the Supreme Court denied certiorari in Thaler v. Perlmutter, confirming that material must have human authorship to receive copyright protection. AI-generated outputs without meaningful human creative input cannot be copyrighted.
What is the Bartz v. Anthropic settlement?
The class action settled for $1.5 billion, with approximately 500,000 copyrighted works receiving roughly $3,000 each. The case established that AI training may be fair use but pirating the source material is not, shifting the legal battleground to how AI companies acquire their training data.
Updates
Corrected an unsupported claim that the DOJ was investigating the Anthropic distillation allegations against DeepSeek, Moonshot AI, and MiniMax (no such investigation has been confirmed); replaced it with the actual corroborated federal response, a proposed bill in Congress and an April 2026 White House national-security memo. Also corrected the Anthropic music-publisher lawsuit's damages figure from a precise "$3.1 billion" to "more than $3 billion," matching the cited source, in four places including the title/meta.
Independently fact-checked against the cited primary sources; governing law re-checked for recent changes
Governing law re-checked for recent changes
Initial publication covering the London Book Fair protest, Anthropic music lawsuit, DeepSeek distillation allegations, and UK copyright policy reversal.
Reviewed and approved by an editor
UK government published its copyright and AI progress report, reversing its preferred opt-out approach after 88% of public consultation respondents opposed it.
The Law Behind This Article
This article rests on the statutory provisions below, held in our own legal record and retrieved from the official source. Tap a section to read the operative text.
United States Code Title 17
§ 107Limitations on exclusive rights: Fair useIn forcecited in 10 of our articles
Notwithstanding the provisions of sections 106 and 106A, the fair use of a copyrighted work, including such use by reproduction in copies or phonorecords or by any other means specified by that section, for purposes such as criticism, comment, news reporting, teaching (including multiple copies for classroom use), scholarship, or research, is not an infringement of copyright. In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include— the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; the nature of the copyrighted work; the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and the effect of the use upon the potential market for or value of the copyrighted work. The fact that a work is unpublished shall not itself bar a finding of fair use if such finding is made upon consideration of all the above factors.
Official text (excerpt) · last checked 2026-07-28 · Read the full text in our law library · Verify at uscode.house.gov
Cited in 883 court opinions in our collectionLatest citing opinion in our collection: 2026
In the courts (editorial summary, independently checked):Harper & Row v. Nation Enterprises (1985) applied the section 107 factors to reject fair use for prepublication quotation from an unpublished memoir, calling market effect the single most important element. Campbell v. Acuff-Rose Music (1994) held commercial character is only one element and rejected a presumption against parody.
Opinions citing this section in our collection:
- Harper & Row, Publishers, Inc. v. Nation Enterprises (Supreme Court of the United States 1985, 471 U.S. 539)✓The Nation, working from a purloined manuscript of Gerald Ford's memoirs, printed about 300 words verbatim and scooped Time's licensed excerpt. The Court held this was not fair use under section 107, weighing the work's unpublished nature and Time's cancellation as market harm.
- Sony Corp. of America v. Universal City Studios, Inc. (Supreme Court of the United States 1984, 464 U.S. 417)✓Universal and Disney sued the maker of the Betamax over home taping of broadcast television. Weighing the section 107 factors, the Court held private noncommercial time-shifting is fair use because the studios showed no likelihood of harm to the market for their works.
- Leadsinger, Inc. v. BMG Music Publishing (Court of Appeals for the Ninth Circuit 2008)✓A karaoke maker sought a declaration that displaying and printing copyrighted lyrics was fair use. Applying the section 107 factors, the Ninth Circuit found the use commercial and not transformative, the lyrics creative and taken whole, and affirmed dismissal of the claim.
Identified automatically from the court opinions citing this section — not a ranking of which case controls.
Also relied on in: How to File a DMCA Takedown on Xvideos (2026 Guide), What Is a DMCA Takedown? Complete Guide, How to File a DMCA Takedown on Etsy (2026 Guide)
Search our full record of US law — 2.1 million sections, every state + federal →
Sources and References
- UK Copyright and AI Progress Report - GOV.UK(gov.uk).gov
- Generative AI and Copyright Law - Congressional Research Service(congress.gov).gov
- 17 U.S.C. 107 - Fair Use(copyright.gov).gov
- Music publishers sue Anthropic for $3B over piracy of 20,000 works - TechCrunch(techcrunch.com)
- Anthropic claims 3 Chinese companies ripped it off - Fortune(fortune.com)
- 10,000 Authors Protest AI With Empty Book - Deadline(deadline.com)
- Authors protest AI at London Book Fair - Euronews(euronews.com)
- UMG Sues Anthropic for $3 Billion - Billboard(billboard.com)
- AI Copyright Cases Update 2026 - Norton Rose Fulbright(nortonrosefulbright.com)
- UK Government Reverses AI Copyright Approach - Lewis Silkin(lewissilkin.com)
- NSTM-4: Memorandum on Adversarial Distillation of American AI Models - White House OSTP (Apr. 23, 2026)(whitehouse.gov).gov
- H.R. 8283 - Deterring American AI Model Theft Act of 2026 - Congress.gov(congress.gov).gov