Chinese Publisher Warns: Books Banned from AI Training, Hard to Enforce but a Stance Must Be Taken

Avatar 0

By NUPIAO News | Reporter: Song Jianan

Recently, sharp-eyed readers noticed something unusual on the copyright pages of books like New Translation of Huangting Jing and Yinfu Jing, published by Huaxia Publishing House. A clear line of warning text now reads: “It is strictly forbidden to use the content of this book for artificial intelligence training. Violators will be held accountable.” The book, annotated by Liu Lianpeng and Gu Baotian, is published and distributed by Huaxia Publishing House Co., Ltd., with the property rights belonging to Sanmin Book Co., Ltd.

On August 4th, according to a report from Red Star News, a representative from Huaxia Publishing House explained that when licensing certain foreign titles, the rights holders often include a clause in the copyright contract prohibiting the use of the book’s text for AI training. “As the party fulfilling the contract, we have an obligation to the rights holders, so we print this statement,” the representative said.

However, the representative acknowledged a harsh reality: if the book’s content is used for AI training without authorization, it’s actually very hard to enforce legal rights. The primary purpose of this printed warning is to make a statement, to remind everyone that books are protected by copyright.

Right now, most traditional print publishers lack the technical means to automatically block the scanning and digitization of their books. Unlike websites, which can use protocols like robots.txt to deter crawlers, physical books have no built-in online defense mechanism. For organized data-collection teams, this warning text is essentially a “gentleman’s agreement.” For ordinary publishers and creators, the practical legal avenues to fight against unauthorized use of their books in AI training remain quite limited.

Legal experts offer some advice: if you discover an infringement, the first step is to thoroughly preserve evidence of ownership, such as the book’s copyright page and physical samples. If you find that a company is mass-scanning books, you can send a formal written request demanding they stop. But the biggest hurdle is gathering proof—it’s nearly impossible for the average person to access the internal training datasets of AI companies. For now, the industry relies heavily on prior agreements and continuously publicizing copyright claims to lower the risk of infringement.

In reality, the copyright battle between the international publishing world and AI companies has been raging for a while. Back in early 2024, AI startup Anthropic launched a secret internal project codenamed “Project Panama,” with the core goal of obtaining “all the books in the world.” In its early days, Anthropic relied on pirated e-book sources like Books3 and shadow libraries to train its models, which led to a wave of class-action lawsuits from authors.

In September of last year, court documents revealed that Anthropic agreed to pay at least $1.5 billion to settle these lawsuits. Under the settlement, Anthropic would pay authors or publishers roughly $3,000 for each of the approximately 500,000 books covered by the agreement.

To dodge the legal risks of piracy, the company pivoted to buying used physical books in bulk to collect high-quality training data. After scanning them, they would literally destroy the originals. Around 2 million paper books were damaged or destroyed in this data-collection process. When the news broke, the “destroying books to train AI” controversy sent shockwaves through the global publishing industry, prompting many publishers to seriously rethink the data security of their physical texts.

Compared to the messy, repetitive, and often factually wrong content found on the open internet, formally published books—having gone through rigorous editing—are coherent and logically sound. They help large language models learn systematic expression and long-form reasoning. Many ancient texts, academic monographs, and out-of-print works don’t have complete digital versions circulating online; the only way to get their full content is by scanning physical copies. A diverse range of published material also expands a model’s knowledge base, balancing skills like formal writing, narrative, and theoretical analysis.

This immense training value makes books an indispensable source of data for large model development. It also explains why some AI companies are willing to go to great lengths and costs to mass-digitize physical books. But the resulting copyright conflicts are now forcing the publishing industry to proactively draw the line.

The academic world is also stepping up. In May of this year, Elsevier, the publishing giant behind journals like The Lancet and Cell, along with four other publishers, jointly sued Meta. They accuse Meta of accessing massive amounts of copyrighted books and journal articles without permission to train its Llama series of models.

According to the lawsuit, in order to get a leg up in the AI arms race, Meta not only used web-scraped datasets containing billions of pages but also downloaded and distributed millions of copyrighted books and paid academic journal articles from pirate sites. Meta is also accused of stripping out copyright notices and author information from these works to cover its tracks.

In response, Meta has argued that using copyrighted content to train AI is “transformative use” and falls under the “fair use” doctrine in U.S. copyright law. The company says it will vigorously defend itself.

Going back a bit further, in March of last year, major French publishers and authors’ associations also sued Meta over the unauthorized mass use of copyrighted content for training its AI models.

To this day, there is still no unified or definitive legal ruling globally. Courts in different jurisdictions have even arrived at completely opposite verdicts in similar cases.

In March of this year, according to a report from 21st Century Business Herald, Tao Kaiyuan, Vice President of the Supreme People’s Court of China, spoke about the challenges generative AI poses to IP judicial rulings. She noted that judges often find it difficult to make decisions by simply “matching” existing laws to the situation, yet they must answer the “judicial questions” raised by the rapid development of AI technology.

In Tao’s view, the key is to strike a balance between the growth of the AI industry and the protection of copyright holders’ legitimate rights. She revealed that the country is currently actively working on drafting judicial policy documents related to AI and data property rights, aiming to provide a clearer legal framework for the application of new technologies.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

Log In / Sign Up

Enter your email to receive a secure code. No password needed.