Is Training AI on Copyrighted Books Legal? The Answer Is Still Evolving

Artificial intelligence models depend on enormous volumes of data. Their training datasets may include books, news articles, academic papers, websites and other published material, often collected without the direct knowledge or consent of the original creators.

For authors and publishers, this raises an important question: can technology companies legally use copyrighted works to build commercial AI systems?

The answer depends on more than whether a protected work was included in a dataset. Courts are also examining how the material was obtained, how it was used and whether the resulting AI product competes with the original creator.

AI training and copyright infringement are not necessarily the same

One of the most influential decisions so far came from a legal dispute involving Anthropic and a group of writers. The court concluded that using books to train a large language model could be considered lawful because the process was transformative.

The reasoning compared AI training with the way a human writer studies existing literature to learn patterns, structures and techniques before creating something new. From this perspective, the model is not necessarily storing works to reproduce them. It is analysing them to develop a different technological capability.

However, the court drew a clear distinction between using copyrighted material and acquiring it illegally. Anthropic was penalised because books used in its training process had been obtained from unauthorised online libraries.

This distinction could become critical for the industry. Even if some forms of AI training qualify as fair use, companies may still face substantial liability when their datasets contain pirated material.

Why fair use matters

Many US copyright disputes involving AI revolve around the principle of fair use. This doctrine permits certain uses of copyrighted content without the owner’s explicit permission, particularly for purposes such as criticism, research, education, parody or other transformative activities.

Courts typically consider several factors:

  • The purpose and character of the use
  • The nature of the original work
  • How much of the work was used
  • The effect on the market for the original content

The final factor may be particularly important in AI cases. If an AI-powered product directly replaces or competes with the source material, courts may be less willing to consider the training process lawful.

Direct competition changes the legal equation

The dispute between Thomson Reuters and Ross Intelligence illustrates this risk. Ross used material from Thomson Reuters to develop an AI-based legal research platform that would compete in the same market.

The court determined that this use was not sufficiently transformative. The new platform had a similar commercial purpose and could affect demand for Thomson Reuters’ original product.

This suggests that courts may treat AI training differently depending on what the resulting system does. A general-purpose model that learns language patterns could receive different legal treatment from a specialised system trained on proprietary content to reproduce a competing service.

Authors may similarly argue that generative AI tools compete with them by producing synthetic books or other written content. So far, however, this argument has not established a definitive legal standard.

Training data and AI-generated content are separate issues

The copyright debate around AI actually involves at least two distinct questions.

The first concerns inputs: can copyrighted material be used to train an AI model?

The second concerns outputs: can content created with AI receive copyright protection?

In Thaler v. Perlmutter, the court concluded that a work created entirely by AI could not be copyrighted because copyright protection requires human authorship.

In practice, the distinction between human and machine authorship is rarely straightforward. A person may use AI to generate ideas, restructure passages, edit drafts or correct language. Even conventional tools such as spell-checking software already contribute to the creative process without becoming authors themselves.

Courts and copyright authorities will therefore need to determine how much human involvement is required before AI-assisted work becomes eligible for protection.

Regulation is struggling to keep pace

Much of the current US copyright framework was developed decades before generative AI. Judges must now apply established legal principles to technologies capable of processing billions or trillions of words and generating new content at scale.

The result is an evolving and sometimes inconsistent legal landscape. Early rulings are already influencing how AI companies source data, document training processes and assess intellectual-property risks. Yet many important lawsuits remain unresolved, and future courts may reach different conclusions.

What this means for businesses

For organisations developing or integrating AI systems, copyright compliance should be treated as part of the technology strategy rather than as a secondary legal concern.

Companies should understand:

  • Where training, fine-tuning and retrieval data comes from
  • Whether the organisation has permission to use that material
  • How datasets and licences are documented
  • Whether the AI product could compete with the owners of its source content
  • How generated outputs are reviewed for potential infringement
  • What contractual protections AI providers offer

The legality of AI training will probably continue to be decided case by case. One principle is already becoming clearer, however: the provenance of data matters.

As AI systems move from experimentation into core business operations, responsible data governance, traceability and intellectual-property controls will become essential parts of building technology that can scale safely.

Source

Control F5 Team
Blog Editor
OUR WORK
Case studies

We have helped 20+ companies in industries like Finance, Transportation, Health, Tourism, Events, Education, Sports.

READY TO DO THIS
Let’s build something together