Part 2 · Chapter 15
Nonexpressive Use As Fair Use
Introduction to Nonexpressive Use
Broadly speaking, nonexpressive use is a concept in copyright law that refers to the use of a copyrighted work in a way that does not communicate the author’s original expression to any new audience. Essentially, it involves copying a work not to be read or enjoyed as a substitute for the original, but rather to extract facts, ideas, or data from it, or to make new independent observations about the work.
Nonexpressive use is particularly relevant in the context of new technologies, such as search engines, text data mining and training machine learning and AI models. When a search engine indexes a website, it makes a copy of the content to analyze and categorize it, but it doesn’t present the full text of the copyrighted work to the user. Researchers and programmers use text data mining (TDM) techniques to analyze large quantities of text to find patterns, trends, and relationships in the data. The original works are copied for this analysis, but the user only sees the resulting data or insights, not the original expression. Machine learning and AI models are often trained on vast datasets that often include copyrighted works. The model “learns” from this data, but the goal is to extract underlying patterns and information to generate new content, not to reproduce the original works in the training set.
Nonexpressive uses are often argued to be fair use. The concept was first articulated by Professor Matthew Sag in Copyright and Copy-Reliant Technology, 103 Northwestern University Law Review 1607 (2009), but has been developed and debated by several legal scholars. See Matthew Sag, Orphan Works as Grist for the Data Mill, 27 Berkeley Technology Law Journal 1503 (2012); Matthew Jockers, Matthew Sag & Jason Schultz, Digital Archives: Don’t Let Copyright Block Data Mining, 490 Nature 29-30 (October 4, 2012). For related and similar arguments, see Ed Lee, Technological Fair Use, 83 Southern California Law Review 797 (2010), Maurizio Borghi & Stavroula Karapapa, Non-Display Uses of Copyright Works, 1 Queen Mary Journal of Intellectual Property 21 (2011), Abraham Drassinower, What’s Wrong With Copying (2015); and more recently, Oren Bracha, The Work of Copyright in the Age of Machine Production (https://ssrn.com/abstract=4581738); and Molly S. Van Houweling, The Freedom to Extract in Copyright Law, 103 North Carolina Law Review 445 (2025).
This chapter begins with an extract of a law review article by Matthew Sag, Fairness and Fair Use in Generative AI, that explains the theory of nonexpressive use and its place in copyright law. The chapter then provides brief summaries of several nonexpressive use cases and a longer extract of the Google Books case. All of these cases are important and reward deeper exploration, but compiling a textbook always entails difficult choices about which materials to present and at what level of detail. Note that Google Books is an important and well written opinion, but it was not a radical departure from the prior caselaw, indeed, after the Second Circuit decision in Authors Guild v. HathiTrust, the decision in Google Books was widely seen as a fait-accompli. The chapter then considers nonexpressive use in the context of AI training as explored in recent cases.
Matthew Sag, Fairness and Fair Use in Generative AI, 92 Fordham Law Review 1887 (2024)
Although we are still a long way from the science fiction version of “artificial general intelligence” that thinks, feels, and refuses to “open the pod bay doors,” recent advances in machine learning and artificial intelligence (AI) have captured the public’s imagination and lawmakers’ interest. We now have large language models (LLMs) that can pass the bar exam, carry on (what passes for) a conversation about almost any topic, create new music, and create new visual art. These artifacts are often indistinguishable from their human-authored counterparts and yet can be produced at a speed and scale surpassing human ability.
Generative AI systems, such as the GPT and LLaMA language models and the Stable Diffusion and Midjourney text-to-image models, were built by ingesting massive quantities of text and images from the Internet. This was done with little or no regard to whether those works were subject to copyright or whether the authors would object to their use.
The rise of generative AI poses important questions for copyright law. These questions, however, are not entirely new. Generative AI gives us yet another context to consider copyright’s most fundamental question: where do the rights of the copyright owner end, and the freedom to use copyrighted works begin? … My aim in this Essay is not to establish that generative AI is, or should be, non-infringing; it is to outline an analytical framework for making that assessment in particular cases.
… What is copyright law about?
The architecture of copyright law orients towards the protection of original expression, not the prohibition of copying. Original expression makes a work copyrightable in the first place, and the contribution of original expression and control over the final form of that expression distinguishes co-authors from mere assistants. Moreover, the exclusive rights of the copyright owner are generally defined by and limited to the communication of original expression to the public. Sometimes, these definitions and limitations are explicit, as with the rights of public performance and public display; sometimes, they are implicit.
The centrality of original expression is most obvious with respect to the distinction that copyright law draws between protectable expression that originates with the author and unprotectable facts, ideas, theories, systems, and methods of operation, whether they spring from the author’s mind or not. Copyright does not forbid the ordinary reader from extracting and reproducing the facts, ideas, or artistic techniques embodied in a work; it encourages her to do so. As Justice O’Connor noted in Feist,
It may seem unfair that much of the fruit of the compiler’s labor may be used by others without compensation. As Justice Brennan has correctly observed, however, this is not ‘some unforeseen byproduct of a statutory scheme.’ It is, rather, ‘the essence of copyright, and a constitutional requirement.” (citing Harper & Row Publishers, Inc. v. Nation Enter., 471 U.S. 539, 589 (1985) (dissenting opinion)).
Moving beyond the idea-expression distinction, copyright’s focus on the communication of original expression to the public is also evident in several other aspects of copyright law: the threshold of substantial similarity (which is determined from the perspective of the ordinary observer and thus inherently a question of how the work is communicated and received); the scope of the publisher’s collective right (which considers how the works are presented to the audience, not the data structure in which they are stored); and the general refusal of courts to base a finding of copyright infringement on unpublished drafts.
In my view, these are not merely isolated examples; they illustrate a more general principle: the copyright owner’s exclusive rights are defined by and limited to the communication of the author’s original expression to the public.
What does this mean for fair use?
This understanding of copyright law gives us a framework to assess claims of fair use and supplies the limiting principle that was missing (or perhaps only implicit) in Judge Leval’s original formulation of transformative use in Towards A Fair Use Standard and the Supreme Court’s adoption thereof in Campbell v. Acuff-Rose Music. Given the centrality of the communication of original expression to the public, the critical function of fair use is to permit uses that, while they may amount to technical acts of copying, do not, in substance, threaten the author’s copyright-protected interest in controlling the communication of their original expression to the public.
Judge Leval was right to focus on transformative use, but the transformative use test would have been far less confusing had it been expressly tied to the benchmark of expressive substitution. Classic fair uses such as parody, commentary, or criticism are not fair use merely because they change the underlying work or convey some new meaning or message. Most movies based on literary works add layers of new meaning and expression, but this does not make them fair use. The movie Rear Window exposed the original expression of a short story called It Had to Be Murder to new audiences and added a lot else besides, but it still required a license. In contrast, 2Live Crew’s consistent parody of Pretty Woman qualified as fair use because the transformations they made were such that the parody posed no risk of expressive substitution to the original.
The Supreme Court’s 2023 decision in Andy Warhol Found. for Visual Arts v. Goldsmith emphasizes that the question of transformative use, or “whether an allegedly infringing use has a further purpose or different character . . . is a matter of degree, and the degree of difference must be weighed against other considerations, like commercialism.” AWF reaffirms the importance of transformative use and implicitly rejects lower court rulings that had found uses to be transformative where there was no significant difference in purpose, merely the addition of a new vibe or aesthetic. Citing Campbell, the majority in AWF explained:
Most copying has some further purpose, in the sense that copying is socially useful ex post. Many secondary works add something new. That alone does not render such uses fair. Rather, the first factor (which is just one factor in a larger analysis) asks ‘whether and to what extent’ the use at issue has a purpose or character different from the original. The larger the difference, the more likely the first factor weighs in favor of fair use. The smaller the difference, the less likely.
AWF helpfully clarifies why a transformative use has featured so prominently in the case law: the more transformative a use is, the less likely it is to substitute for the copyright owner’s original expression. Using the author’s work to reflect on the original is an intrinsically different purpose; that difference in purpose makes expressive substitution unlikely. In contrast, merely adding an overlay of new expression while leaving the original expression intact provides no such comfort. The majority in AWF rightly focuses on whether the defendant’s use is likely to substitute for the author’s original expression and makes that the measure of when the defendant’s use is sufficiently transformative.
Non-Expressive Use
This brings us to what I call the “non-expressive use” cases. United States courts have consistently held that technical acts of copying that do not communicate an author’s original expression to a new audience are fair use. Examples of non-expressive uses include copying object code to extract uncopyrightable facts and interoperability keys (“reverse engineering”), an automated process of copying student term papers to compare to other papers for plagiarism detection, copying HTML webpages to make a search engine index, copying printed library books to allow researchers to conduct statistical analysis of the contents of whole collections of books, and copying printed library books to create a search engine index.
The case law indicates that even though these non-expressive uses involved significant amounts of copying, they did not interfere with the interest in original expression that copyright is designed to protect. Each use involved copying as an intermediate step towards producing something that either did not contain the original expression of the underlying work or contained a trivial amount. Courts have consistently held that non-expressive uses (although not labeled as such) are fair use.
In a 1992 decision, Sega Enterprises, Ltd. v. Accolade, Inc., 977 F.2d 1510 (9th Cir. 1992) and again in 2000 in Sony Computer Entertainment, Inc. v. Connectix Corp., 203 F.3d 596 (9th Cir. 2000) the Ninth Circuit held that reverse engineering object code—a process that involves making several copies to extract vital but uncopyrightable elements needed to create interoperable programs—was fair use. In Sega, the court referred to copying to extract uncopyrightable elements as “a legitimate, essentially non-exploitative purpose . . . .” In Sony Computer Entertainment, the court expressly recognized that “the fair use doctrine preserves public access to the ideas and functional elements embedded in copyrighted computer software programs.”
In A.V. ex rel. Vanderhye v. iParadigms, LLC, 562 F.3d 630 (4th Cir. 2009) the Fourth Circuit held that copying student papers into a reference database for comparison against new student papers was fair use. In the 2014 case Authors Guild, Inc. v. HathiTrust, 755 F.3d 87 (2d Cir. 2014) the Second Circuit held that making digital versions of printed library books for research purposes that included text data mining and machine learning was fair use. A differently constituted panel of the Second Circuit reached much the same conclusion in 2015 in Authors Guild v. Google, Inc., 804 F.3d 202 (2d Cir. 2015) In Google Books, the court addressed both the complete copying of millions of library books to make them searchable and the display of small snippets of the books in search result menus. The complete copying of books is an example of non-expressive use; the snippet displays illustrate the application of a more traditional transformative use analysis.
When courts have declined to find fair use in cases that are superficially similar to those discussed above, it is invariably because the challenged use was not non-expressive, and thus, on the facts presented, the potential substitution effect was too significant. For example, in Associated Press v. Meltwater U.S. Holdings, Inc., 931 F. Supp. 2d 537 (S.D.N.Y. 2013) the U.S. District Court for the Southern District of New York held that fair use did not justify the actions of a media monitoring company, Meltwater. Meltwater scraped news articles on the web to provide its subscribers with excerpts and analytics. However, the lawsuit did not challenge Meltwater’s use of copyrighted news articles to provide metadata and analytics to its subscribers, even though these services also necessitated copying. The court noted that this was “an entirely separate service” and implied that, if challenged, it would be transformative and thus fair use. Instead of attacking Meltwater’s non-expressive use, the Associated Press focused on the length and significance of Meltwater’s extracts provided to subscribers. The court agreed that Meltwater’s extracts were too long and too close to the heart of the work; it also held that Meltwater had failed to show that the amount of the extracts was reasonable in light of its stated purpose to operate like a search engine.
In a similar fashion, in Fox News Network, LLC v. TVEyes, Inc., 883 F.3d 169 (2d Cir. 2018) the Second Circuit held that a media monitoring service that copied and electronically searched television broadcasts went beyond the scope of fair use when it allowed users to save, watch, and share ten-minute long video clips of the copyrighted programs. In the court’s view, those ten-minute video clips would “likely provide TVEyes’s users with all of the Fox programming that they seek and the entirety of the message conveyed by Fox to authorized viewers of the original.” In other words, the court was concerned that rather than primarily providing information about the content of particular news segments, the length of the video clips was such that they would substitute for those segments in their entirety. The district court in TVEyes held that copying for search alone was fair use, and Fox did not contest this ruling on appeal.
The concept of non-expressive use explains these cases and differentiates them from more amorphous calls for “fair learning” that are difficult to reconcile with the case law. For example, American Geophysical Union v. Texaco, Inc. 60 F.3d 913(2d Cir.1994) would be a prime candidate for “fair learning,” but it was not a non-expressive use. In that case, researchers employed by Texaco had a practice of photocopying scientific articles for later reading, which the majority of the Second Circuit thought went beyond fair use because the articles were copied “for the same basic purpose that one would normally seek to obtain the original—to have it available on his shelf for ready reference if and when [the researcher] needed to look at it” and because the publishers had established the prospect of “a workable market for institutional users to obtain licenses for the right to produce their own copies of individual articles via photocopying.”
My point up until now has been: (1) that there are general principles internal to copyright that courts can look at to understand the function and application of the fair use doctrine; (2) that one such principle is that copyright was never intended to convey sole and despotic dominion over every use of every word—copyright exists, by and large, to prevent the communication of the author’s original expression to new audiences without authorization or compensation; (3) that this realization suggests a positive vision for fair use—the critical function of fair use is to permit uses that, although they may amount to technical acts of copying, do not in substance threaten the author’s interest in controlling the communication of their original expression to the public; and (4) that non-expressive uses meet this threshold. Non-expressive uses, by definition, pose no threat of direct expressive substitution. Admittedly, non-expressive uses generate information about works: although such information has value and utility and may even influence the demand for the original work, metadata and other uncopyrightable abstract concepts do not satisfy the public’s appetite for the author’s original expression.
[The article then addresses how this framework to generative AI]
Notes and questions
(1) Why does the author argue for a principle of nonexpressive use, but reject “fair learning”? Aren’t these the same thing? For more, see Mark A. Lemley & Bryan Casey, Fair Learning, 99 Texas Law Review 743 (2021) (arguing that we should treat “fair learning as a lawful purpose under the first factor.”)
The Nonexpressive Use Cases
Sega Enterprises Ltd. v. Accolade, Inc., 977 F.2d 1510 (9th Cir. 1992)
In Sega v. Accolade the Ninth Circuit addressed whether reverse engineering a copyrighted computer program to access its unprotected functional elements constitutes fair use. Software is typically distributed in object code, a machine-readable format, and though protected by copyright, it contains unprotectable elements like ideas and functional interfaces. Accolade reverse-engineered Sega’s code not to exploit it directly but to develop its own Genesis-compatible games. The court found that because disassembly was the only way to access the functional requirements for compatibility, and because Accolade used the code for a legitimate, non-exploitative purpose, this intermediate copying qualified as fair use.
The court’s analysis emphasized the distinction between protected expression and unprotected ideas or functional components. It held that enabling interoperability and fostering competition serve the public interest and align with copyright’s constitutional goal of promoting progress in the useful arts. While Sega experienced some indirect market effects from Accolade’s entry, the court found that allowing such competition better served the statutory objectives than protecting Sega’s market control. Accordingly, the court concluded that all relevant fair use factors favored Accolade and that its reverse engineering of Sega’s code was lawful, though Sega remained free to challenge the final games on separate grounds.
Sony Computer Entertainment v. Connectix, 203 F.3d 596 (9th Cir. 2000)
Almost a decade after Sega, the Ninth Circuit addressed the same issue in Sony Computer Entertainment v. Connectix. From the beginning of its decision, the court emphasized the importance of the idea expression distinction:
[W]e are called upon once again to apply the principles of copyright law to computers and their software, to determine what must be protected as expression and what must be made accessible to the public as function.
Consistent with its decision in Sega, the court held that intermediate copying of software could be protected as fair use if the copying was necessary to gain access to the functional elements of the software. The court expressly recognized (at 603) that “the fair use doctrine preserves public access to the ideas and functional elements embedded in copyrighted computer software programs.”
Since Sega v. Accolade courts have consistently held that making unauthorized copies of a computer program, as a necessary step in reverse engineering, is fair use. This outcome was implicitly confirmed by the fact that Congress chose to include circumventing encryption for the purpose of reverse engineering as an allowable exception to the anti-circumvention provisions of the DMCA. See Section 1201(f) of the Copyright Act.
Assessment Technologies of Wis., LLC v. WIREdata, 350 F.3d 640 (7th Cir. 2003)
In WIREdata, the plaintiff, Assessment Technologies (AT), sued WIREdata for copyright infringement and theft of trade secrets. AT had developed and copyrighted a software program called “Market Drive,” which was used by tax assessors to compile property data collected for tax purposes. WIREdata, which provides information to real estate brokers, sought to obtain this raw, public-domain data from municipalities that were AT’s licensees. AT sued to prevent WIREdata from accessing the data, arguing that it couldn’t be extracted from the Market Drive program without infringing on its copyright. The district court sided with AT and issued an injunction, but the Seventh Circuit Court of Appeals reversed this decision.
The appeals court found that while AT’s compilation format (the structure of its database) was original enough to be copyrighted, the underlying data was not created by AT, was not copyrightable, and was in the public domain. Applying the reverse engineering cases discussed above, the court held that extracting this non-copyrightable data from the database, even if it required copying the program in the process, was a fair use. The court does not say much about fair use, seemingly because the issue was so clear cut: “AT would lose this copyright case even if the raw data were so entangled with Market Drive that they could not be extracted without making a copy of the program.” It was also noted that AT was attempting to use its copyright to “sequester uncopyrightable data,” which could constitute copyright misuse. The court concluded that WIREdata could obtain the data through various methods without infringing on AT’s copyright, and therefore, the injunction was improper.
Lexmark Intern. v. Static Control Components, 387 F.3d 522 (6th Cir. 2004)
In 2004, the Sixth Circuit Court of Appeals heard a case between printer manufacturer Lexmark and Static Control Components (SCC), a company that sells components for remanufactured toner cartridges. The facts of the case centered on two computer programs owned by Lexmark: the Toner Loading Program, which is on the toner cartridge and measures toner levels, and the much larger Printer Engine Program, which is in the printer itself and controls various printer functions. To prevent third-party reuse of its discounted “Prebate” cartridges, Lexmark used an authentication sequence to ensure that only a microchip with an identical copy of the Toner Loading Program could function with its printers. SCC developed a competing microchip that mimicked Lexmark’s, enabling remanufactured cartridges to work with Lexmark printers. SCC’s chip contained an exact copy of the Toner Loading Program, which it claimed was necessary for compatibility. The district court granted a preliminary injunction against SCC, but the Court of Appeals vacated the injunction.
This case is extracted earlier in this book as it pertains to the merger doctrine. The majority found the Toner Loading Program was not copyrightable, but it still considered how fair use might apply. The majority explained that while a commercial purpose may generally weigh against fair use, SCC’s use was not meant to exploit Lexmark’s creative expression. Instead, SCC used the program for a different, non-infringing purpose: to make its cartridges interoperable with Lexmark printers: “Under these circumstances, it is far from clear that SCC copied the Toner Loading Program for its commercial value as a copyrighted work.” (emphasis original). The court also determined that the fourth fair use factor—the effect on the potential market—was misconstrued by the lower court. The correct analysis should have focused on the market for the copyrighted work itself (the Toner Loading Program), not the market for Lexmark’s toner cartridges. The court found no evidence that an independent market existed for a program as simple as the Toner Loading Program. Therefore, SCC’s use of the program did not diminish the value of a potential market that was not shown to exist, and this factor should have weighed in favor of SCC.
A.V. ex rel. Vanderhye v. iParadigms, LLC, 562 F.3d 630 (4th Cir. 2009)
In A.V. ex rel. Vanderhye v. iParadigms, LLC, the Fourth Circuit held that it was fair use for Turnitin, a commercial plagiarism detection service, to archive student papers submitted through its platform. Although the plaintiffs, high school students, alleged copyright infringement, the court found that Turnitin’s use of the papers was highly transformative. The system stored the papers solely to compare future submissions for signs of plagiarism—a purpose entirely distinct from the original educational intent of the student work. Even though iParadigms profited from its service, the court emphasized that the transformative nature of the use outweighed its commercial character, especially given the public benefit of academic integrity.
The court also rejected arguments that the use was not transformative because Turnitin stored the papers unaltered or because the system was imperfect at catching plagiarism. A transformative use need not alter the original work, nor must it achieve its goal flawlessly. The function of Turnitin—to detect textual similarities as a means of preventing academic dishonesty—was found to be different enough in purpose from the original expression to qualify as fair use. Since iParadigms did not exploit the expressive content of the papers and no meaningful market for high school essays was harmed, the court affirmed that all fair use factors favored the defendant.
Authors Guild v. HathiTrust, 755 F.3d 87 (2d Cir. 2014)
HathiTrust is extracted at length in the chapter on fair use in particular contexts, and the discussion there should be read alongside this one. In brief: HathiTrust is a repository founded in 2008 by a group of major universities to store and preserve millions of digitized works, most of them scanned by Google as part of the Library Project. Its functions include full-text search across the whole corpus, access for users with certified print disabilities, and support for text mining and scholarly research.
The Second Circuit held that the full-text search function was “quintessentially transformative”: it provided a new utility that did not substitute for reading the books, and making complete digital copies was necessary to deliver it. The court found no cognizable market harm, rejecting arguments built on hypothetical licensing markets and speculative data-breach risk. Digitization for search, and access for the print-disabled, were fair use. The court remanded the question whether copying for preservation was fair use.
For present purposes the case matters because it is the first appellate decision to treat mass digitization for search as transformative — the proposition Google Books took up a year later, and the one the generative AI cases now test.
Authors Guild v. Google, Inc., 804 F.3d 202 (2d Cir. 2015)
Circuit Judge Leval
This copyright dispute tests the boundaries of fair use. Plaintiffs, who are authors of published books under copyright, sued Google, Inc. (“Google”) for copyright infringement in the United States District Court for the Southern District of New York (Chin, J.). They appeal from the grant of summary judgment in Google’s favor. Through its Library Project and its Google Books project, acting without permission of rights holders, Google has made digital copies of tens of millions of books, including Plaintiffs’, that were submitted to it for that purpose by major libraries. Google has scanned the digital copies and established a publicly available search function. An Internet user can use this function to search without charge to determine whether the book contains a specified word or term and also see “snippets” of text containing the searched-for terms. In addition, Google has allowed the participating libraries to download and retain digital copies of the books they submit, under agreements which commit the libraries not to use their digital copies in violation of the copyright laws. These activities of Google are alleged to constitute infringement of Plaintiffs’ copyrights. Plaintiffs sought injunctive and declaratory relief as well as damages.
Google defended on the ground that its actions constitute “fair use,” which, under 17 U.S.C. § 107, is “not an infringement.” The district court agreed. Authors Guild, Inc. v. Google Inc., 954 F.Supp.2d 282, 294 (S.D.N.Y.2013). Plaintiffs brought this appeal.
Google Books and the Google Library Project
Google’s Library Project, which began in 2004, involves bi-lateral agreements between Google and a number of the world’s major research libraries. Under these agreements, the participating libraries select books from their collections to submit to Google for inclusion in the project. Google makes a digital scan of each book, extracts a machine-readable text, and creates an index of the machine-readable text of each book. Google retains the original scanned image of each book, in part so as to improve the accuracy of the machine-readable texts and indices as image-to-text conversion technologies improve.
Since 2004, Google has scanned, rendered machine-readable, and indexed more than 20 million books, including both copyrighted works and works in the public domain. The vast majority of the books are non-fiction, and most are out of print. All of the digital information created by Google in the process is stored on servers protected by the same security systems Google uses to shield its own confidential information.
The digital corpus created by the scanning of these millions of books enables the Google Books search engine. Members of the public who access the Google Books website can enter search words or terms of their own choice, receiving in response a list of all books in the database in which those terms appear, as well as the number of times the term appears in each book. A brief description of each book, entitled “About the Book,” gives some rudimentary additional information, including a list of the words and terms that appear with most frequency in the book. It sometimes provides links to buy the book online and identifies libraries where the book can be found. The search tool permits a researcher to identify those books, out of millions, that do, as well as those that do not, use the terms selected by the researcher. Google notes that this identifying information instantaneously supplied would otherwise not be obtainable in lifetimes of searching.
No advertising is displayed to a user of the search function. Nor does Google receive payment by reason of the searcher’s use of Google’s link to purchase the book.
The search engine also makes possible new forms of research, known as “text mining” and “data mining.” Google’s “ngrams” research tool draws on the Google Library Project corpus to furnish statistical information to Internet users about the frequency of word and phrase usage over centuries. This tool permits users to discern fluctuations of interest in a particular subject over time and space by showing increases and decreases in the frequency of reference and usage in different periods and different linguistic regions. It also allows researchers to comb over the tens of millions of books Google has scanned in order to examine “word frequencies, syntactic patterns, and thematic markers” and to derive information on how nomenclature, linguistic usage, and literary style have changed over time. Authors Guild, Inc., 954 F.Supp.2d at 287. The district court gave as an example “tracking the frequency of references to the United States as a single entity (‘the United States is’) versus references to the United States in the plural (‘the United States are’) and how that usage has changed over time.”
[Judge Leval is referring to an example in the amicus brief filed on behalf of Digital Humanities Researchers. The illustration below is not in the judgement.]
United States “is” vs “are” illustration

The Google Books search function also allows the user a limited viewing of text. In addition to telling the number of times the word or term selected by the searcher appears in the book, the search function will display a maximum of three “snippets” containing it. A snippet is a horizontal segment comprising ordinarily an eighth of a page. Each page of a conventionally formatted book in the Google Books database is divided into eight non-overlapping horizontal segments, each such horizontal segment being a snippet. (Thus, for such a book with 24 lines to a page, each snippet is comprised of three lines of text.) Each search for a particular word or term within a book will reveal the same three snippets, regardless of the number of computers from which the search is launched. Only the first usage of the term on a given page is displayed. Thus, if the top snippet of a page contains two (or more) words for which the user searches, and Google’s program is fixed to reveal that particular snippet in response to a search for either term, the second search will duplicate the snippet already revealed by the first search, rather than moving to reveal a different snippet containing the word because the first snippet was already revealed. Google’s program does not allow a searcher to increase the number of snippets revealed by repeated entry of the same search term or by entering searches from different computers. A searcher can view more than three snippets of a book by entering additional searches for different terms. However, Google makes permanently unavailable for snippet view one snippet on each page and one complete page out of every ten — a process Google calls “blacklisting.”
Google also disables snippet view entirely for types of books for which a single snippet is likely to satisfy the searcher’s present need for the book, such as dictionaries, cookbooks, and books of short poems. Finally, since 2005, Google will exclude any book altogether from snippet view at the request of the rights holder by the submission of an online form.
Under its contracts with the participating libraries, Google allows each library to download copies — of both the digital image and machine-readable versions — of the books that library submitted to Google for scanning (but not of books submitted by other libraries). The agreements between Google and the libraries, although not in all respects uniform, require the libraries to abide by copyright law in utilizing the digital copies they download and to take precautions to prevent dissemination of their digital copies to the public at large. Participant libraries have downloaded at least 2.7 million digital copies of their own volumes.
[The Law of Fair Use]
[The court restated the fair use framework: the § 107 preamble and four factors, Campbell’s emphasis on transformative purpose under factor one, and the Harper & Row observation that factor four is “undoubtedly the single most important element.” That framework is set out at length in the fair use chapters above.]
The Search and Snippet View Functions
A. Factor One
(1) Transformative purpose. Campbell’s explanation of the first factor’s inquiry into the “purpose and character” of the secondary use focuses on whether the new work, “in Justice Story’s words, merely supersedes the objects’ of the original creation, or instead adds something new, with a further purpose. It asks, in other words, whether and to what extent the new work is transformative.” 510 US at 578-579 (citations omitted).1 While recognizing that a transformative use is “not absolutely necessary for a finding of fair use,” the opinion further explains that the “goal of copyright, to promote science and the arts, is generally furthered by the creation of transformative works” and that “such works thus lie at the heart of the fair use doctrine’s guarantee of breathing space within the confines of copyright.” Id. at 579. In other words, transformative uses tend to favor a fair use finding because a transformative use is one that communicates something new and different from the original or expands its utility, thus serving copyright’s overall objective of contributing to public knowledge.
The word “transformative” cannot be taken too literally as a sufficient key to understanding the elements of fair use. It is rather a suggestive symbol for a complex thought, and does not mean that any and all changes made to an author’s original text will necessarily support a finding of fair use. The Supreme Court’s discussion in Campbell gave important guidance on assessing when a transformative use tends to support a conclusion of fair use. The defendant in that case defended on the ground that its work was a parody of the original and that parody is a time-honored category of fair use. Explaining why parody makes a stronger, or in any event more obvious, claim of fair use than satire, the Court stated,
[T]he heart of any parodist’s claim to quote from existing material ... is the use of ... a prior author’s composition to ... comment[] on that author’s works.... If, on the contrary, the commentary has no critical bearing on the substance or style of the original composition, which the alleged infringer merely uses to get attention or to avoid the drudgery in working up something fresh, the claim to fairness in borrowing from another’s work diminishes accordingly (if it does not vanish).... Parody needs to mimic an original to make its point, and so has some claim to use the creation of its victim’s ... imagination, whereas satire can stand on its own two feet and so requires justification for the very act of borrowing.
Id. at 580-81 (emphasis added). In other words, the would-be fair user of another’s work must have justification for the taking. A secondary author is not necessarily at liberty to make wholesale takings of the original author’s expression merely because of how well the original author’s expression would convey the secondary author’s different message. Among the best recognized justifications for copying from another’s work is to provide comment on it or criticism of it. A taking from another author’s work for the purpose of making points that have no bearing on the original may well be fair use, but the taker would need to show a justification. This part of the Supreme Court’s discussion is significant in assessing Google’s claim of fair use because, as discussed extensively below, Google’s claim of transformative purpose for copying from the works of others is to provide otherwise unavailable information about the originals.
A further complication that can result from oversimplified reliance on whether the copying involves transformation is that the word “transform” also plays a role in defining “derivative works,” over which the original rights holder retains exclusive control. Section 106 of the Act specifies the exclusive right of the copyright owner “(2) to prepare derivative works based upon the copyrighted work.” See 17 U.S.C. § 106. The statute defines derivative works largely by example, rather than explanation. The examples include “translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgement, condensation,” to which list the statute adds “any other form in which a work may be ... transformed.” 17 U.S.C. § 101 (emphasis added). As we noted in Authors Guild, Inc. v. HathiTrust, “paradigmatic examples of derivative works include the translation of a novel into another language, the adaptation of a novel into a movie or play, or the recasting of a novel as an e-book or an audiobook.” 755 F.3d 87, 95 (2d Cir.2014). While such changes can be described as transformations, they do not involve the kind of transformative purpose that favors a fair use finding. The statutory definition suggests that derivative works generally involve transformations in the nature of changes of form. 17 U.S.C. § 101. By contrast, copying from an original for the purpose of criticism or commentary on the original or provision of information about it, tends most clearly to satisfy Campbell’s notion of the “transformative” purpose involved in the analysis of Factor One.18
Footnote 18: The Seventh Circuit takes the position that the kind of secondary use that favors satisfaction of the fair use test is better described as a “complementary” use, referring to how a hammer and nail complement one another in that together they achieve results that neither can accomplish on its own. Ty, Inc. v. Publications International, Ltd., 292 F.3d 512, 517-518 (7th Cir.2002); see also Kienitz v. Sconnie Nation LLC, 766 F.3d 756, 758 (7th Cir.2014). We do not find the term “complementary” particularly helpful in explaining fair use. The term would encompass changes of form that are generally understood to produce derivative works, rather than fair uses, and, at the same time, would fail to encompass copying for purposes that are generally and properly viewed as creating fair uses. When a novel is converted into film, for example, the original novel and the film ideally complement one another in that each contributes to achieving results that neither can accomplish on its own. The invention of the original author combines with the cinematographic interpretive skills of the filmmaker to produce something that neither could have produced independently. Nonetheless, at least when the intention of the film is to make a “motion picture version” of the novel, 17 U.S.C. § 101, without undertaking to parody it or to comment on it, the film is generally understood to be a derivative work, which under § 106, falls within the exclusive rights of the copyright owner. Although they complement one another, the film is not a fair use. At the same time, when a secondary work quotes an original for the purpose of parodying it, or discrediting it by exposing its inaccuracies, illogic, or dishonesty, such an undertaking is not within the exclusive prerogatives of the rights holder; it produces a fair use. Yet, when the purpose of the second is essentially to destroy the first, the two are not comfortably described as complementaries that combine to produce together something that neither could have produced independently of the other. We recognize, as just noted above, that the word “transformative,” if interpreted too broadly, can also seem to authorize copying that should fall within the scope of an author’s derivative rights. Attempts to find a circumspect shorthand for a complex concept are best understood as suggestive of a general direction, rather than as definitive descriptions.
With these considerations in mind, we first consider whether Google’s search and snippet views functions satisfy the first fair use factor with respect to Plaintiffs’ rights in their books. (The question whether these functions might infringe upon Plaintiffs’ derivative rights is discussed in the next Part.)
(2) Search Function. We have no difficulty concluding that Google’s making of a digital copy of Plaintiffs’ books for the purpose of enabling a search for identification of books containing a term of interest to the searcher involves a highly transformative purpose, in the sense intended by Campbell. Our court’s exemplary discussion in HathiTrust informs our ruling. That case involved a dispute that is closely related, although not identical, to this one. Authors brought claims of copyright infringement against HathiTrust, an entity formed by libraries participating in the Google Library Project to pool the digital copies of their books created for them by Google. The suit challenged various usages HathiTrust made of the digital copies. Among the challenged uses was HathiTrust’s offer to its patrons of “full-text searches,” which, very much like the search offered by Google Books to Internet users, permitted patrons of the libraries to locate in which of the digitized books specific words or phrases appeared. 755 F.3d at 98. (HathiTrust’s search facility did not include the snippet view function, or any other display of text.) We concluded that both the making of the digital copies and the use of those copies to offer the search tool were fair uses. Id. at 105.
Notwithstanding that the libraries had downloaded and stored complete digital copies of entire books, we noted that such copying was essential to permit searchers to identify and locate the books in which words or phrases of interest to them appeared. Id. at 97. We concluded “that the creation of a full-text searchable database is a quintessentially transformative use ... [as] the result of a word search is different in purpose, character, expression, meaning, and message from the page (and the book) from which it is drawn.” Id. We cited A.V. ex rel. Vanderhye v. iParadigms, LLC, 562 F.3d 630, 639-40 (4th Cir.2009), Perfect 10, Inc. v. Amazon.com, Inc., 508 F.3d 1146, 1165 (9th Cir.2007), and Kelly v. Arriba Soft Corp., 336 F.3d 811, 819 (9th Cir.2003) as examples of cases in which courts had similarly found the creation of complete digital copies of copyrighted works to be transformative fair uses when the copies “served a different function from the original.” HathiTrust, 755 F.3d at 97.
As with HathiTrust (and iParadigms), the purpose of Google’s copying of the original copyrighted books is to make available significant information about those books, permitting a searcher to identify those that contain a word or term of interest, as well as those that do not include reference to it. In addition, through the ngrams tool, Google allows readers to learn the frequency of usage of selected words in the aggregate corpus of published books in different historical periods. We have no doubt that the purpose of this copying is the sort of transformative purpose described in Campbell as strongly favoring satisfaction of the first factor.
We recognize that our case differs from HathiTrust in two potentially significant respects. First, HathiTrust did not “display to the user any text from the underlying copyrighted work,” 755 F.3d at 91, whereas Google Books provides the searcher with snippets containing the word that is the subject of the search. Second, HathiTrust was a nonprofit educational entity, while Google is a profit-motivated commercial corporation. We discuss those differences below.
(3) Snippet View. Plaintiffs correctly point out that this case is significantly different from HathiTrust in that the Google Books search function allows searchers to read snippets from the book searched, whereas HathiTrust did not allow searchers to view any part of the book. Snippet view adds important value to the basic transformative search function, which tells only whether and how often the searched term appears in the book. Merely knowing that a term of interest appears in a book does not necessarily tell the searcher whether she needs to obtain the book, because it does not reveal whether the term is discussed in a manner or context falling within the scope of the searcher’s interest. For example, a searcher seeking books that explore Einstein’s theories, who finds that a particular book includes 39 usages of “Einstein,” will nonetheless conclude she can skip that book if the snippets reveal that the book speaks of “Einstein” because that is the name of the author’s cat. In contrast, the snippet will tell the searcher that this is a book she needs to obtain if the snippet shows that the author is engaging with Einstein’s theories.
Google’s division of the page into tiny snippets is designed to show the searcher just enough context surrounding the searched term to help her evaluate whether the book falls within the scope of her interest (without revealing so much as to threaten the author’s copyright interests). Snippet view thus adds importantly to the highly transformative purpose of identifying books of interest to the searcher. With respect to the first factor test, it favors a finding of fair use (unless the value of its transformative purpose is overcome by its providing text in a manner that offers a competing substitute for Plaintiffs’ books, which we discuss under factors three and four below).
(4) Google’s Commercial Motivation. Plaintiffs also contend that Google’s commercial motivation weighs in their favor under the first factor. Google’s commercial motivation distinguishes this case from HathiTrust, as the defendant in that case was a non-profit entity founded by, and acting as the representative of, libraries. Although Google has no revenues flowing directly from its operation of the Google Books functions, Plaintiffs stress that Google is profit-motivated and seeks to use its dominance of book search to fortify its overall dominance of the Internet search market, and that thereby Google indirectly reaps profits from the Google Books functions.
For these arguments Plaintiffs rely primarily on two sources. First is Congress’s specification in spelling out the first fair use factor in the text of § 107 that consideration of the “purpose and character of the [secondary] use” should “include whether such use is of a commercial nature or is for nonprofit educational purposes.” Second is the Supreme Court’s assertion in dictum in Sony Corporation of America v. Universal City Studios that “every commercial use of copyrighted material is presumptively ... unfair.” 464 US 417, 451 (1984). If that were the extent of precedential authority on the relevance of commercial motivation, Plaintiffs’ arguments would muster impressive support. However, while the commercial motivation of the secondary use can undoubtedly weigh against a finding of fair use in some circumstances, the Supreme Court, our court, and others have eventually recognized that the Sony dictum was enormously overstated.19
Footnote 19: Campbell, 510 US at 583-84; Cariou v. Prince, 714 F.3d 694, 708 (2d Cir.2013); Castle Rock Entm’t, Inc. v. Carol Pub. Grp., Inc., 150 F.3d 132, 141-42 (2d Cir.1998); Perfect 10, Inc. v. Amazon.com, Inc., 508 F.3d 1146, 1165 (9th Cir.2007); Kelly v. Arriba Soft Corp., 336 F.3d 811, 819 (9th Cir.2003); see also Monge v. Maya Magazines, Inc., 688 F.3d 1164, 1172 (9th Cir.2012) (noting that Campbell “debunked the notion that Sony called for a ‘hard evidentiary presumption’ that commercial use is presumptively unfair.”)
The Sixth Circuit took the Sony dictum at its word in Acuff-Rose Music, Inc. v. Campbell, concluding that, because the defendant rap music group’s spoof of the plaintiff’s ballad was done for profit, it could not be fair use. 972 F.2d 1429, 1436-1437 (6th Cir.1992). The Supreme Court reversed on this very point, observing that “Congress could not have intended” such a broad presumption against commercial fair uses, as “nearly all of the illustrative uses listed in the preamble paragraph of § 107 ... are generally conducted for profit in this country.” Campbell, 510 US at 584 (internal quotation marks and citations omitted). The Court emphasized Congress’s statement in the House Report to the effect that the commercial or nonprofit character of a work is “not conclusive” but merely “a fact to be ‘weighed along with other[s] in fair use decisions.’” Id. at 585. In explaining the first fair use factor, the Court clarified that “the more transformative the [secondary] work, the less will be the significance of other factors, like commercialism, that may weigh against a finding of fair use.” Id. at 579.
Our court has since repeatedly rejected the contention that commercial motivation should outweigh a convincing transformative purpose and absence of significant substitutive competition with the original. See Cariou v. Prince, 714 F.3d 694, 708 (2d Cir.2013); Castle Rock Entertainment, Inc. v. Carol Publication Group, Inc., 150 F.3d 132, 141-42 (2d Cir.1998).
While we recognize that in some circumstances, a commercial motivation on the part of the secondary user will weigh against her, especially, as the Supreme Court suggested, when a persuasive transformative purpose is lacking, Campbell, 510 US at 579, we see no reason in this case why Google’s overall profit motivation should prevail as a reason for denying fair use over its highly convincing transformative purpose, together with the absence of significant substitutive competition, as reasons for granting fair use. Many of the most universally accepted forms of fair use, such as news reporting and commentary, quotation in historical or analytic books, reviews of books, and performances, as well as parody, are all normally done commercially for profit.20
Footnote 20: Just as there is no reason for presuming that a commercial use is not a fair use, which would defeat the most widely accepted and logically justified areas of fair use, there is likewise no reason to presume categorically that a nonprofit educational purpose should qualify as a fair use. Authors who write for educational purposes, and publishers who invest substantial funds to publish educational materials, would lose the ability to earn revenues if users were permitted to copy the materials freely merely because such copying was in the service of a nonprofit educational mission. The publication of educational materials would be substantially curtailed if such publications could be freely copied for non-profit educational purposes.
B. Factor Two
The second fair use factor directs consideration of the “nature of the copyrighted work.” While the “transformative purpose” inquiry discussed above is conventionally treated as a part of first factor analysis, it inevitably involves the second factor as well. One cannot assess whether the copying work has an objective that differs from the original without considering both works, and their respective objectives.
The second factor has rarely played a significant role in the determination of a fair use dispute. The Supreme Court in Harper & Row made a passing observation in dictum that, “[t]he law generally recognizes a greater need to disseminate factual works than works of fiction or fantasy.” 471 US 539, 563 (1985). Courts have sometimes speculated that this might mean that a finding of fair use is more favored when the copying is of factual works than when copying is from works of fiction. However, while the copyright does not protect facts or ideas set forth in a work, it does protect that author’s manner of expressing those facts and ideas. At least unless a persuasive fair use justification is involved, authors of factual works, like authors of fiction, should be entitled to copyright protection of their protected expression. The mere fact that the original is a factual work therefore should not imply that others may freely copy it. Those who report the news undoubtedly create factual works. It cannot seriously be argued that, for that reason, others may freely copy and re-disseminate news reports.21
Footnote 21: We think it unlikely that the Supreme Court meant in its concise dictum that secondary authors are at liberty to copy extensively from the protected expression of the original author merely because the material is factual. What the Harper & Row dictum may well have meant is that, because in the case of factual writings, there is often occasion to test the accuracy of, to rely on, or to repeat their factual propositions, and such testing and reliance may reasonably require quotation (lest a change of expression unwittingly alter the facts), factual works often present well justified fair uses, even if the mere fact that the work is factual does not necessarily justify copying of its protected expression.
In considering the second factor in HathiTrust, we concluded that it was “not dispositive,” 755 F.3d at 98, commenting that courts have hardly ever found that the second factor in isolation played a large role in explaining a fair use decision. The same is true here. While each of the three Plaintiffs’ books in this case is factual, we do not consider that as a boost to Google’s claim of fair use. If one (or all) of the plaintiff works were fiction, we do not think that would change in any way our appraisal. Nothing in this case influences us one way or the other with respect to the second factor considered in isolation. To the extent that the “nature” of the original copyrighted work necessarily combines with the “purpose and character” of the secondary work to permit assessment of whether the secondary work uses the original in a “transformative” manner, as the term is used in Campbell, the second factor favors fair use not because Plaintiffs’ works are factual, but because the secondary use transformatively provides valuable information about the original, rather than replicating protected expression in a manner that provides a meaningful substitute for the original.
C. Factor Three
The third statutory factor instructs us to consider “the amount and substantiality of the portion used in relation to the copyrighted work as a whole.” The clear implication of the third factor is that a finding of fair use is more likely when small amounts, or less important passages, are copied than when the copying is extensive, or encompasses the most important parts of the original. The obvious reason for this lies in the relationship between the third and the fourth factors. The larger the amount, or the more important the part, of the original that is copied, the greater the likelihood that the secondary work might serve as an effectively competing substitute for the original, and might therefore diminish the original rights holder’s sales and profits.
(1) Search Function. The Google Books program has made a digital copy of the entirety of each of Plaintiffs’ books. Notwithstanding the reasonable implication of Factor Three that fair use is more likely to be favored by the copying of smaller, rather than larger, portions of the original, courts have rejected any categorical rule that a copying of the entirety cannot be a fair use. Complete unchanged copying has repeatedly been found justified as fair use when the copying was reasonably appropriate to achieve the copier’s transformative purpose and was done in such a manner that it did not offer a competing substitute for the original.24
Footnote 24: See cases cited supra note 17; see also Bill Graham Archives v. Dorling Kindersley Ltd., 448 F.3d 605, 613 (2d Cir.2006) (Copying the entirety of a work is sometimes necessary to make a fair use of the work).
The Supreme Court said in Campbell that “the extent of permissible copying varies with the purpose and character of the use” and characterized the relevant questions as whether “the amount and substantiality of the portion used ... are reasonable in relation to the purpose of the copying,” Campbell, 510 US at 586-587, noting that the answer to that question will be affected by “the degree to which the [copying work] may serve as a market substitute for the original or potentially licensed derivatives,” id. at 587-588 (finding that, in the case of a parodic song, “how much ... is reasonable will depend, say, on the extent to which the song’s overriding purpose and character is to parody the original or, in contrast, the likelihood that the parody may serve as a market substitute for the original”).
In HathiTrust, our court concluded in its discussion of the third factor that “because it was reasonably necessary for the [HathiTrust Digital Library] to make use of the entirety of the works in order to enable the full-text search function, we do not believe the copying was excessive.” 755 F.3d at 98. As with HathiTrust, not only is the copying of the totality of the original reasonably appropriate to Google’s transformative purpose, it is literally necessary to achieve that purpose. If Google copied less than the totality of the originals, its search function could not advise searchers reliably whether their searched term appears in a book (or how many times).
While Google makes an unauthorized digital copy of the entire book, it does not reveal that digital copy to the public. The copy is made to enable the search functions to reveal limited, important information about the books. With respect to the search function, Google satisfies the third factor test, as illuminated by the Supreme Court in Campbell.
(2) Snippet View. Google’s provision of snippet view makes our third factor inquiry different from that inquiry in HathiTrust. What matters in such cases is not so much “the amount and substantiality of the portion used” in making a copy, but rather the amount and substantiality of what is thereby made accessible to a public for which it may serve as a competing substitute. In HathiTrust, notwithstanding the defendant’s full-text copying, the search function revealed virtually nothing of the text of the originals to the public. Here, through the snippet view, more is revealed to searchers than in HathiTrust.
Without doubt, enabling searchers to see portions of the copied texts could have determinative effect on the fair use analysis. The larger the quantity of the copyrighted text the searcher can see and the more control the searcher can exercise over what part of the text she sees, the greater the likelihood that those revelations could serve her as an effective, free substitute for the purchase of the plaintiff’s book. We nonetheless conclude that, at least as presently structured by Google, the snippet view does not reveal matter that offers the marketplace a significantly competing substitute for the copyrighted work.
Google has constructed the snippet feature in a manner that substantially protects against its serving as an effectively competing substitute for Plaintiffs’ books. In the Background section of this opinion, we describe a variety of limitations Google imposes on the snippet function. These include the small size of the snippets (normally one eighth of a page), the blacklisting of one snippet per page and of one page in every ten, the fact that no more than three snippets are shown — and no more than one per page — for each term searched, and the fact that the same snippets are shown for a searched term no matter how many times, or from how many different computers, the term is searched. In addition, Google does not provide snippet view for types of books, such as dictionaries and cookbooks, for which viewing a small segment is likely to satisfy the searcher’s need. The result of these restrictions is, so far as the record demonstrates, that a searcher cannot succeed, even after long extended effort to multiply what can be revealed, in revealing through a snippet search what could usefully serve as a competing substitute for the original.
The blacklisting, which permanently blocks about 22% of a book’s text from snippet view, is by no means the most important of the obstacles Google has designed. While it is true that the blacklisting of 22% leaves 78% of a book theoretically accessible to a searcher, it does not follow that any large part of that 78% is in fact accessible. The other restrictions built into the program work together to ensure that, even after protracted effort over a substantial period of time, only small and randomly scattered portions of a book will be accessible. In an effort to show what large portions of text searchers can read through persistently augmented snippet searches, Plaintiffs’ counsel employed researchers over a period of weeks to do multiple word searches on Plaintiffs’ books. In no case were they able to access as much as 16% of the text, and the snippets collected were usually not sequential but scattered randomly throughout the book. Because Google’s snippets are arbitrarily and uniformly divided by lines of text, and not by complete sentences, paragraphs, or any measure dictated by content, a searcher would have great difficulty constructing a search so as to provide any extensive information about the book’s use of that term. As snippet view never reveals more than one snippet per page in response to repeated searches for the same term, it is at least difficult, and often impossible, for a searcher to gain access to more than a single snippet’s worth of an extended, continuous discussion of the term.
The fact that Plaintiffs’ searchers managed to reveal nearly 16% of the text of Plaintiffs’ books overstates the degree to which snippet view can provide a meaningful substitute. At least as important as the percentage of words of a book that are revealed is the manner and order in which they are revealed. Even if the search function revealed 100% of the words of the copyrighted book, this would be of little substitutive value if the words were revealed in alphabetical order, or any order other than the order they follow in the original book. It cannot be said that a revelation is “substantial” in the sense intended by the statute’s third factor if the revelation is in a form that communicates little of the sense of the original. The fragmentary and scattered nature of the snippets revealed, even after a determined, assiduous, time-consuming search, results in a revelation that is not “substantial,” even if it includes an aggregate 16% of the text of the book. If snippet view could be used to reveal a coherent block amounting to 16% of a book, that would raise a very different question beyond the scope of our inquiry.
D. Factor Four
The fourth fair use factor, “the effect of the [copying] use upon the potential market for or value of the copyrighted work,” focuses on whether the copy brings to the marketplace a competing substitute for the original, or its derivative, so as to deprive the rights holder of significant revenues because of the likelihood that potential purchasers may opt to acquire the copy in preference to the original. Because copyright is a commercial doctrine whose objective is to stimulate creativity among potential authors by enabling them to earn money from their creations, the fourth factor is of great importance in making a fair use assessment. See Harper & Row, 471 US at 566 (describing the fourth factor as “undoubtedly the single most important element of fair use”).
Campbell stressed the close linkage between the first and fourth factors, in that the more the copying is done to achieve a purpose that differs from the purpose of the original, the less likely it is that the copy will serve as a satisfactory substitute for the original. 510 US at 591. Consistent with that observation, the HathiTrust court found that the fourth factor favored the defendant and supported a finding of fair use because the ability to search the text of the book to determine whether it includes selected words “does not serve as a substitute for the books that are being searched.” 755 F.3d at 100.
However, Campbell’s observation as to the likelihood of a secondary use serving as an effective substitute goes only so far. Even if the purpose of the copying is for a valuably transformative purpose, such copying might nonetheless harm the value of the copyrighted original if done in a manner that results in widespread revelation of sufficiently significant portions of the original as to make available a significantly competing substitute. The question for us is whether snippet view, notwithstanding its transformative purpose, does that. We conclude that, at least as snippet view is presently constructed, it does not.
Especially in view of the fact that the normal purchase price of a book is relatively low in relation to the cost of manpower needed to secure an arbitrary assortment of randomly scattered snippets, we conclude that the snippet function does not give searchers access to effectively competing substitutes. Snippet view, at best and after a large commitment of manpower, produces discontinuous, tiny fragments, amounting in the aggregate to no more than 16% of a book. This does not threaten the rights holders with any significant harm to the value of their copyrights or diminish their harvest of copyright revenue.
We recognize that the snippet function can cause some loss of sales. There are surely instances in which a searcher’s need for access to a text will be satisfied by the snippet view, resulting in either the loss of a sale to that searcher, or reduction of demand on libraries for that title, which might have resulted in libraries purchasing additional copies. But the possibility, or even the probability or certainty, of some loss of sales does not suffice to make the copy an effectively competing substitute that would tilt the weighty fourth factor in favor of the rights holder in the original. There must be a meaningful or significant effect “upon the potential market for or value of the copyrighted work.” 17 U.S.C. § 107(4).
Furthermore, the type of loss of sale envisioned above will generally occur in relation to interests that are not protected by the copyright. A snippet’s capacity to satisfy a searcher’s need for access to a copyrighted book will at times be because the snippet conveys a historical fact that the searcher needs to ascertain. For example, a student writing a paper on Franklin D. Roosevelt might need to learn the year Roosevelt was stricken with polio. By entering “Roosevelt polio” in a Google Books search, the student would be taken to (among numerous sites) a snippet from page 31 of Richard Thayer Goldberg’s The Making of Franklin D. Roosevelt (1981), telling that the polio attack occurred in 1921. This would satisfy the searcher’s need for the book, eliminating any need to purchase it or acquire it from a library. But what the searcher derived from the snippet was a historical fact. Author Goldberg’s copyright does not extend to the facts communicated by his book. It protects only the author’s manner of expression. Hoehling v. Universal City Studios, Inc., 618 F.2d 972, 974 (2d Cir.1980) (“A grant of copyright in a published work secures for its author a limited monopoly over the expression it contains.”) (emphasis added). Google would be entitled, without infringement of Goldberg’s copyright, to answer the student’s query about the year Roosevelt was afflicted, taking the information from Goldberg’s book. The fact that, in the case of the student’s snippet search, the information came embedded in three lines of Goldberg’s writing, which were superfluous to the searcher’s needs, would not change the taking of an unprotected fact into a copyright infringement.
Even if the snippet reveals some authorial expression, because of the brevity of a single snippet and the cumbersome, disjointed, and incomplete nature of the aggregation of snippets made available through snippet view, we think it would be a rare case in which the searcher’s interest in the protected aspect of the author’s work would be satisfied by what is available from snippet view, and rarer still — because of the cumbersome, disjointed, and incomplete nature of the aggregation of snippets made available through snippet view — that snippet view could provide a significant substitute for the purchase of the author’s book.
Accordingly, considering the four fair use factors in light of the goals of copyright, we conclude that Google’s making of a complete digital copy of Plaintiffs’ works for the purpose of providing the public with its search and snippet view functions (at least as snippet view is presently designed) is a fair use and does not infringe Plaintiffs’ copyrights in their books.
III. Derivative Rights in Search and Snippet View
Plaintiffs next contend that, under Section 106(2), they have a derivative right in the application of search and snippet view functions to their works, and that Google has usurped their exclusive market for such derivatives.
There is no merit to this argument. As explained above, Google does not infringe Plaintiffs’ copyright in their works by making digital copies of them, where the copies are used to enable the public to get information about the works, such as whether, and how often they use specified words or terms (together with peripheral snippets of text, sufficient to show the context in which the word is used but too small to provide a meaningful substitute for the work’s copyrighted expression). The copyright resulting from the Plaintiffs’ authorship of their works does not include an exclusive right to furnish the kind of information about the works that Google’s programs provide to the public. For substantially the same reasons, the copyright that protects Plaintiffs’ works does not include an exclusive derivative right to supply such information through query of a digitized copy.
The extension of copyright protection beyond the copying of the work in its original form to cover also the copying of a derivative reflects a clear and logical policy choice. An author’s right to control and profit from the dissemination of her work ought not to be evaded by conversion of the work into a different form. The author of a book written in English should be entitled to control also the dissemination of the same book translated into other languages, or a conversion of the book into a film. The copyright of a composer of a symphony or song should cover also conversions of the piece into scores for different instrumentation, as well as into recordings of performances.
This policy is reflected in the statutory definition, which explains the scope of the “derivative” largely by examples — including “a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgement, [or] condensation” — before adding, “or any other form in which a work may be recast, transformed, or adapted.” 17 U.S.C. § 101. As noted above, this definition, while imprecise, strongly implies that derivative works over which the author of the original enjoys exclusive rights ordinarily are those that re-present the protected aspects of the original work, i.e., its expressive content, converted into an altered form, such as the conversion of a novel into a film, the translation of a writing into a different language, the reproduction of a painting in the form of a poster or post card, recreation of a cartoon character in the form of a three-dimensional plush toy, adaptation of a musical composition for different instruments, or other similar conversions. If Plaintiffs’ claim were based on Google’s converting their books into a digitized form and making that digitized version accessible to the public, their claim would be strong. But as noted above, Google safeguards from public view the digitized copies it makes and allows access only to the extent of permitting the public to search for the very limited information accessible through the search function and snippet view. The program does not allow access in any substantial way to a book’s expressive content. Nothing in the statutory definition of a derivative work, or of the logic that underlies it, suggests that the author of an original work enjoys an exclusive derivative right to supply information about that work of the sort communicated by Google’s search functions.
Plaintiffs seek to support their derivative claim by a showing that there exist, or would have existed, paid licensing markets in digitized works, such as those provided by the Copyright Clearance Center or the previous, revenue-generating version of the Google Partners Program. Plaintiffs also point to the proposed settlement agreement rejected by the district court in this case, according to which Google would have paid authors for its use of digitized copies of their works. The existence or potential existence of such paid licensing schemes does not support Plaintiffs’ derivative argument. The access to the expressive content of the original that is or would have been provided by the paid licensing arrangements Plaintiffs cite is far more extensive than that which Google’s search and snippet view functions provide. Those arrangements allow or would have allowed public users to read substantial portions of the book. Such access would most likely constitute copyright infringement if not licensed by the rights holders. Accordingly, such arrangements have no bearing on Google’s present programs, which, in a non-infringing manner, allow the public to obtain limited data about the contents of the book, without allowing any substantial reading of its text.
Plaintiffs also seek to support their derivative claim by a showing that there is a current unpaid market in licenses for partial viewing of digitized books, such as the licenses that publishers currently grant to the Google Partners program and Amazon’s Search Inside the Book program to display substantial portions of their books. Plaintiffs rely on Infinity Broadcast Corporation v. Kirkwood, 150 F.3d 104 (2nd Cir.1998) and United States v. American Society of Composers, Authors and Publishers (ASCAP), 599 F.Supp.2d 415 (S.D.N.Y.2009) for the proposition that “a secondary use that replaces a comparable service licensed by the copyright holder, even without charge, may cause market harm.” In the cases cited, however, the purpose of the challenged secondary uses was not the dissemination of information about the original works, which falls outside the protection of the copyright, but was rather the re-transmission, or re-dissemination, of their expressive content. Those precedents do not support the proposition Plaintiffs assert — namely that the availability of licenses for providing unprotected information about a copyrighted work, or supplying unprotected services related to it, gives the copyright holder the right to exclude others from providing such information or services.
While the telephone ringtones at issue in the ASCAP case Plaintiffs cite are superficially comparable to Google’s snippets in that both consist of brief segments of the copyrighted work, in a more significant way they are fundamentally different. While it is true that Google’s snippets display a fragment of expressive content, the fragments it displays result from the appearance of the term selected by the searcher in an otherwise arbitrarily selected snippet of text. Unlike the reading experience that the Google Partners program or the Amazon Search Inside the Book program provides, the snippet function does not provide searchers with any meaningful experience of the expressive content of the book. Its purpose is not to communicate copyrighted expression, but rather, by revealing to the searcher a tiny segment surrounding the searched term, to give some minimal contextual information to help the searcher learn whether the book’s use of that term will be of interest to her. The segments taken from copyrighted music as ringtones, in contrast, are selected precisely because they play the most famous, beloved passages of the particular piece — the expressive content that members of the public want to hear when their phone rings. The value of the ringtone to the purchaser is not that it provides information but that it provides a mini-performance of the most appealing segment of the author’s expressive content. There is no reason to think the courts in the cited cases would have come to the same conclusion if the service being provided by the secondary user had been simply to identify to a subscriber in what key a selected composition was written, the year it was written, or the name of the composer. These cases, and the existence of unpaid licensing schemes for substantial viewing of digitized works, do not support Plaintiffs’ derivative works argument.
[IV. Risks of Hacking, and V. Distribution to Participant Libraries]
[The court disposed of two further arguments. First, the plaintiffs said that Google’s storage of their digitized books exposed them to the risk of hacking, so that a security breach would put the full texts into circulation. The court accepted that a breach would be a serious matter, but found the record showed Google’s security practices to be careful and the risk speculative; a theoretical vulnerability does not defeat a fair use that is otherwise established. Second, Google had returned a digital copy of each scanned book to the library that supplied it. The court held this lawful because the libraries were themselves entitled to make such copies for the non-infringing purposes identified in HathiTrust, and Google, acting as their agent, could do for them what they could do for themselves. Nothing in the record suggested the libraries would use the copies to infringe.]
Notes and questions
(1) In both Authors Guild, Inc. v. HathiTrust, 755 F.3d 87 (2d Cir. 2014) and Authors Guild v. Google, Inc., 804 F.3d 202 (2d Cir. 2015), the Second Circuit held that library digitization for search related purposes was transformative and ultimately fair use. How does the defendants’ use in these cases compare to Campbell v. Acuff-Rose?
(2) Notice that in Google Books, the court addressed both the complete copying of millions of library books to make them searchable, and the display of small snippets of the books in search result menus. The complete copying is an example of non-expressive use; the snippet displays illustrate the application of a more traditional transformative use analysis. The court held that the display of three-line snippets to add context to book search results was transformative in purpose and that it was reasonable in proportion to that purpose. Those snippets allowed a user to verify that a book suggested by the search engine was in fact relevant to her interests. In addition, the snippets were so brief that they did not pose any risk of fulfilling the readers’ demand for the original expression of the underlying manuscripts.
(3) There are other cases that are consistent with a “right to extract” uncopyrighted elements. For more, see Molly S. Van Houweling, The Freedom to Extract in Copyright Law, 103 North Carolina Law Review 445 (2025). The merger analysis in Lexmark Intern. v. Static Control Components, 387 F.3d 522 (6th Cir. 2004) would be a good example of how this framing is broader than nonexpressive use.
The Generative AI Copyright Cases
Obviously, interest in how copyright deals with nonexpressive uses has increased dramatically with the advent of generative AI. The nonexpressive use framing is not universally accepted, and there is a healthy academic debate about it. We return to the criticisms after the cases, when you will be in a position to judge them; the literature is collected at the end of the chapter.
The debate over nonexpressive use is far from the only important issue in the generative AI copyright cases, but it is a good place to start.
Is AI-training fair use?
Thomson Reuters v. Ross Intelligence 765 F.Supp.3d 382 (2025)
In Thomson Reuters v. Ross Intelligence the U.S. District Court for the District of Delaware granted summary judgment to Thomson Reuters, finding that Ross Intelligence’s use of copyrighted content from the Westlaw legal research platform to train its AI-powered legal research tool was not a fair use and constituted copyright infringement.
The Delaware court found that Ross’s use was not “transformative.” It emphasized that Ross used the Westlaw headnotes to create a directly competing product, which had a purpose and character similar to the original work. The Ross Intelligence case is significant as one of the first major judicial decisions on the application of copyright law to AI training. Whether the court’s reasoning would apply to generative AI is uncertain. Recognizing that the case involved “controlling questions of law as to which there is substantial ground for difference of opinion,” the court certified two key questions for an interlocutory appeal to the Third Circuit: (1) whether the Westlaw headnotes and Key Number System are sufficiently original for copyright protection, and (2) whether Ross’s use of the headnotes was a fair use.
The Third Circuit heard argument in Philadelphia on 11 June 2026, before Judges Restrepo, Montgomery-Reeves and Bove. The questioning concentrated on two of the four factors: the first, and whether Ross’s purpose in using the headnotes was transformative, and the fourth, and whether Thomson Reuters had shown harm to an actual or potential market for licensing headnotes to train legal research tools. No decision has yet issued. When it comes, it will in all likelihood be the first appellate ruling on fair use in AI training, although the court may well try to distinguish the technology at issue here from generative AI.
Bartz v. Anthropic PBC, 2025 WL 1741691 (N.D. Cal. June 23, 2025) (first extract)
This extract focuses on the core fair use issue relating to LLM training, a later extract will address other issues in this case.
William Alsup, United States District Judge
INTRODUCTION
Defendant Anthropic PBC is an AI software firm founded by former OpenAI employees in January 2021. Its core offering is an AI software service called Claude. When a user prompts Claude with text, Claude quickly responds with text — mimicking human reading and writing. Claude can do so because Anthropic trained Claude — or rather trained large language models or LLMs underlying various versions of Claude — using books and other texts selected from a central library Anthropic had assembled.
In August 2024, the three individual authors brought this putative class action complaining that Anthropic had infringed its federal copyrights by pirating copies for its library and by reproducing them to train its LLMs. Anthropic now moves for summary judgment on fair use only. Fair use is a legal question for the judge with underlying fact questions, if any, for the jury. To prevail on summary judgment, Anthropic must rely on undisputed facts and/or factual inferences favoring the opposing side. Anthropic thus bears the burdens of production and persuasion in this motion.
ANALYSIS
1. The Purpose and Character of the Use.
For a given use at issue, the first factor addresses “the purpose and character of [that] use, including whether [it] is of a commercial nature or is for nonprofit educational purposes.” 17 U.S.C. § 107(1).
A. The Copies Used to Train Specific LLMs.
All agree that one use at issue was training LLMs to receive text inputs and return text outputs. More specifically, Anthropic used copies of Authors’ copyrighted works to iteratively map statistical relationships between every text-fragment and every sequence of text-fragments so that a completed LLM could receive new text inputs and return new text outputs as if it were a human reading prompts and writing responses. Authors further argue — and this order takes for granted — that such training entailed “memoriz[ing]” works by “compress[ing]” copies of those works into the LLM. The LLMs “memorize[d] A LOT, like A LOT”. Regardless, the “purpose and character” of using works to train LLMs was transformative — spectacularly so.
To repeat and be clear: Authors do not allege that any LLM output provided to users infringed upon Authors’ works. Our record shows the opposite. Users interacted only with the Claude service, which placed additional software between the user and the underlying LLM to ensure that no infringing output ever reached the users. This was akin to the limits Google imposed on how many snippets of text from any one book could be seen by any one user through its Google Books service, preventing its search tool from devolving into a reading tool. Google, 804 F.3d at 222. Here, if the outputs seen by users had been infringing, Authors would have a different case. And, if the outputs were ever to become infringing, Authors could bring such a case. But that is not this case.
Instead, Authors challenge only the inputs, not the outputs, of these LLMs. They point to the fully trained LLMs and the Claude service only to shed light on how training itself uses copies of their works and the ways the Claude service could be used to produce still other works that would compete with their works. This order does the same. Authors’ arguments that the training use is not transformative are unavailing.
First, Authors argue that using works to train Claude’s underlying LLMs was like using works to train any person to read and write, so Authors should be able to exclude Anthropic from this use. But Authors cannot rightly exclude anyone from using their works for training or learning as such. Everyone reads texts, too, then writes new texts. They may need to pay for getting their hands on a text in the first instance. But to make anyone pay specifically for the use of a book each time they read it, each time they recall it from memory, each time they later draw upon it when writing new things in new ways would be unthinkable. For centuries, we have read and re-read books. We have admired, memorized, and internalized their sweeping themes, their substantive points, and their stylistic solutions to recurring writing problems.
Second, to that last point, Authors further argue that the training was intended to memorize their works’ creative elements — not just their works’ non-protectable ones. But this is the same argument. Again, Anthropic’s LLMs have not reproduced to the public a given work’s creative elements, nor even one author’s identifiable expressive style (assuming arguendo that these are even copyrightable). Yes, Claude has outputted grammar, composition, and style that the underlying LLM distilled from thousands of works. But if someone were to read all the modern-day classics because of their exceptional expression, memorize them, and then emulate a blend of their best writing, would that violate the Copyright Act? Of course not. Copyright does not extend to “method[s] of operation, concept[s], [or] principle[s]” “illustrated[ ] or embodied in [a] work.” 17 U.S.C. § 102(b).
Third, Authors next argue that computers nonetheless should not be allowed to do what people do.
Authors cite a decision seeming to say as much. But the judge there twice emphasized while discussing “purpose and character” of the use that what was trained was “not generative AI (AI that writes new content itself).” Rather, what was trained — using a proprietary system for finding court opinions in response to a given legal topic — was a competing AI tool for finding court opinions in response to a given legal topic. That was not transformative. Thomson Reuters Enter. Centre GmbH v. Ross Intell. Inc., 765 F. Supp. 3d 382, 398 (D. Del. 2025).
A better analogue to our facts would be an AI tool trained — using court opinions, and briefs, law review articles, and the like — to receive legal prompts and respond with fresh legal writing. And, on facts much like those, a different court came out the other way. It found fair use. White v. W. Pub. Corp., 29 F. Supp. 3d 396, 400 (S.D.N.Y. 2014).
The latter use stood sufficiently “orthogonal” to anything that any copyright owner rightly could expect to control. See Warhol, 598 U.S. at 538–40. It could thus be freed up for the copyist to use, promoting the progress of science and the arts, without diminishing the incentive to create.
In short, the purpose and character of using copyrighted works to train LLMs to generate new text was quintessentially transformative. Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different. If this training process reasonably required making copies within the LLM or otherwise, those copies were engaged in a transformative use.
The first factor favors fair use for the training copies.
2. The Nature of the Copyrighted Work.
The second fair use factor is “the nature of the copyrighted work.” 17 U.S.C. § 107(2). This factor “calls for recognition that some works are closer to the core of intended copyright protection than others, with the consequence that fair use is more difficult to establish when the former works are copied.” Campbell, 510 U.S. at 586. For one thing, less protection is due published works than unpublished ones. For another, less protection is due “factual works than works of fiction or fantasy.” Harper & Row, 471 U.S. at 563. But less protection is not no protection. Even the arrangement of otherwise unprotectable facts surpasses the low bar for a protectable original work of authorship. Google, 804 F.3d at 220.
Here, Anthropic accepts that all of Authors’ books — all published, whether non-fiction or fiction — contained expressive elements. And, as set out above, this order accepts Authors’ view of the evidence that their works were chosen for their expressive qualities in building a central library and then in training specific LLMs.
The main function of the second factor is to help assess the other factors: to reveal differences between the nature of the works at issue and the nature of their secondary use (above), and to reveal any relation between the amount and substantiality of each work taken and the secondary use (next). E.g., Campbell, 510 U.S. at 586; Kelly, 336 F.3d at 820; Google, 804 F.3d at 220; HathiTrust, 755 F.3d at 98; Bill Graham Archives v. Dorling Kindersley Ltd., 448 F.3d 605, 612–13 (2d Cir. 2006).
The second factor points against fair use for all copies alike.
3. The Amount and Substantiality of the Portion Used.
The third fair use factor is “the amount and substantiality of the portion” of the copyrighted work used by the accused. 17 U.S.C. § 107(3). The crux of this factor is whether the amount was “reasonable in relation to the purpose of the copying.” Campbell, 510 U.S. at 586. Thus, the amount of copying is considered first against the work itself, then more importantly against the proposed transformative purpose. See Warhol, 598 U.S. at 543 & n.18.
A. The Copies Used to Train Specific LLMs.
Copies selected for inclusion in training sets were selected because they were complete and because they contained rich protectible expression, or so this order accepts the record shows for Authors. Was all this copying reasonably necessary to the transformative use?
Yes.
“What matters [ ] is not so much ‘the amount and substantiality of the portion used’ in making a copy, but rather the amount and substantiality of what is thereby made accessible to a public [in the purported secondary use] for which it may serve as a competing substitute [for the primary use].” Google, 804 F.3d at 222. Here, once again, there is no allegation of any traceable connection between the Claude service’s outputs and Authors’ works. The copying used to train the LLMs underlying Claude was thus especially reasonable.
In response, Authors object primarily that the copying used in training was both extremely extensive and not strictly necessary.
As to extensive copying, it is true that entire works were copied. And, copying entire works militates against a finding of fair use. But we just addressed why Authors’ argument is misdirected. The copies that count for this factor are those that would merely serve the same use as the work’s ordinary one. Authors do not allege such copying. The accused use here of the incremental copies is as orthogonal as can be imagined to the ordinary use of a book.
As to strict necessity, Authors make a stronger point. When a productive use is made possible only by borrowing from a specific work, fair use climbs towards its zenith. When a productive use is possible without that borrowing, fair use falls to its nadir — and the borrowing deserves a particularly compelling justification. See Warhol, 598 U.S. at 543 & n.18, 547. Here, it is true that Anthropic could have used some other books or no books at all for training its LLMs — or so this order accepts the record shows for Authors. But Anthropic has presented a compelling explanation for why it was reasonably necessary to use them anyway.
For one thing, all agree Anthropic needed billions of words to train any given LLM. If using only books, Anthropic would have needed millions of books per model. If using a set comprising only a small fraction of books and a larger fraction of other texts, Anthropic still would have needed hundreds of thousands of books. Authors contend that because Anthropic showed it could use such smaller sets of books, it surely could have used no books at all — or at least not their books. But Authors forget that “reasonably necessary” does not mean “strictly necessary.” Authors do not contest that the volume of text required to train an LLM is monumental. Because using so many works was reasonably necessary, using any one work for actually training LLMs was about as reasonable as the next.
For another thing, no output to the public was even alleged to be infringing. So, yes, Authors’ works were chosen as the strongest examples of writing. But the compelling benefits of training the LLMs on strong examples were not offset by revelations to the public of any portion of the works themselves. What was copied was therefore especially reasonable and compelling.
The third factor thus favors fair use for the training copies.
4. The Effect of the Use Upon the Market for or Value of the Copyrighted Work.
The final factor is “the effect of the use upon the potential market for or value of the copyrighted work.” 17 U.S.C. § 107(4). This factor points against fair use when a copyist makes copies available that displace demand for copies the copyright owner already makes available or readily could. Texaco, 60 F.3d at 926–28 (reproduced copies); Dr. Seuss Enters., L.P. v. ComicMix LLC, 983 F.3d 443, 461 (9th Cir. 2020) (derivative copies). “While the first factor considers whether and to what extent an original work and secondary use [in principle could] have substitutable purposes, the fourth factor focuses on actual or potential market substitution.” Warhol, 598 U.S. at 536 n.12 (emphasis added).
A. The Copies Used to Train Specific LLMs.
The copies used to train specific LLMs did not and will not displace demand for copies of Authors’ works, or not in the way that counts under the Copyright Act.
Again, Authors concede that training LLMs did not result in any exact copies nor even infringing knockoffs of their works being provided to the public. If that were not so, this would be a different case. Authors remain free to bring that case in the future should such facts develop.
Instead, Authors contend generically that training LLMs will result in an explosion of works competing with their works — such as by creating alternative summaries of factual events, alternative examples of compelling writing about fictional events, and so on. This order assumes that is so. But Authors’ complaint is no different than it would be if they complained that training schoolchildren to write well would result in an explosion of competing works. This is not the kind of competitive or creative displacement that concerns the Copyright Act. The Act seeks to advance original works of authorship, not to protect authors against competition. Sega, 977 F.2d at 1523–24.
Authors next contend that training LLMs displaced (or will) an emerging market for licensing their works for the narrow purpose of training LLMs. Anthropic argues that transactional costs would exceed Anthropic’s expected benefit from any such bargain, prompting it to cease dealing with any rightsholders or else to cease developing such technology altogether. Our record could support either account — so this order must assume Authors are correct. A market could develop. Even so, such a market for that use is not one the Copyright Act entitles Authors to exploit.
None of the cases cited by Authors requires a different result. All contemplated losses of something the Copyright Act properly protected — not the kinds of fair uses for which a copyright owner cannot rightly expect to control. See TVEyes, Inc., 883 F.3d at 181 (use of a right legally reserved to and factually already being licensed by copyright owner); Texaco, 60 F.3d at 931 (same); Ringgold v. BET, Inc., 126 F.3d 70, 80–81 (2d Cir. 1997) (use of a right legally reserved to and factually likely to be marketable by copyright owner — displaying images of her artistic work in television shows); cf. Seltzer v. Green Day, Inc., 725 F.3d 1170, 1179 (9th Cir. 2013) (no evidence use could be or “was likely to” be marketable).
The fourth factor thus favors fair use for the training copies.
5. Overall Analysis.
After the four factors and any others deemed relevant are “explored, [ ] the results [are] weighed together, in light of the purposes of copyright.” Campbell, 510 U.S. at 578.
The copies used to train specific LLMs were justified as a fair use. Every factor but the nature of the copyrighted work favors this result. The technology at issue was among the most transformative many of us will see in our lifetimes.
This order grants summary judgment for Anthropic that the training use was a fair use.
Notes and questions
(1) Judge Alsup did not say precisely nonexpressive uses are fair uses, but is his decision consistent with that approach? Why did the court conclude that copying books to train an LLM was highly transformative and how did that conclusion affect his analysis of the other fair use factors?
(2) As suggested by Sag in Copyright Safety for Generative AI and Henderson et al in Foundation Models and Fair Use, the fair use status of generative AI is likely to depend on a number of considerations including the degree of memorization, the effectiveness of “copyright safety” measures intended to make inadvertent memorization unavailable. How did Anthropic’s copyright safety measures make its fair use argument stronger?
(3) Nonexpressive uses and other highly transformative uses are not always fair use. Sag argues that nonexpressive uses are generally fair uses, but in part of Fairness and Fair Use in Generative AI not extracted above, he also acknowledges that even when the defendant wins the first fair use factor, there may still be cases where the fourth factor weighs against fair use. Keep this in mind as you read the next case.
(4) Notice how much work is being done by the premise that no output infringed. Judge Alsup takes it as given that “Authors do not allege that any LLM output provided to users infringed upon Authors’ works,” and the analysis that follows is built on that footing. That premise is exactly what is contested in In re OpenAI, Inc. Copyright Infringement Litigation, No. 1:25-md-03143 (S.D.N.Y.), the consolidated proceeding against OpenAI. On 27 October 2025 Judge Sidney Stein denied OpenAI’s motion to dismiss direct infringement claims premised on ChatGPT’s outputs, applying the “more discerning observer” test and holding that a reasonable factfinder could find some outputs — summaries of George R.R. Martin’s novels — substantially similar to the novels themselves. The ruling expressly did not reach fair use, and summary judgment briefing in the MDL runs to November 2026, so the fair use question there remains open. This is a very strange decision. For more, see Matthew Sag, Copyright Winter is Coming (to Wikipedia?), 6 November 2025, https://matthewsag.com/copyright-winter-is-coming-to-wikipedia/:
On October 27, 2025, Judge Sidney Stein of the Southern District of New York denied OpenAI’s motion to dismiss claims that ChatGPT outputs infringed the rights of authors such as George R.R. Martin and David Baldacci. The opinion suggests that short summaries of popular works of fiction are very likely infringing (unless fair use comes to the rescue).
This is a fundamental assault on the idea, expression, distinction as applied to works of fiction. It places thousands of Wikipedia entries in the copyright crosshairs and suggests that any kind of summary or analysis of a work of fiction is presumptively infringing.
Kadrey v. Meta Platforms, Inc., 2025 WL 1752484 (N.D. Cal. June 25, 2025)
VINCE CHHABRIA, United States District Judge
Companies are presently racing to develop generative artificial intelligence models—software products that are capable of generating text, images, videos, or sound based on materials they’ve previously been “trained” on. Because the performance of a generative AI model depends on the amount and quality of data it absorbs as part of its training, companies have been unable to resist the temptation to feed copyright-protected materials into their models—without getting permission from the copyright holders or paying them for the right to use their works for this purpose. This case presents the question whether such conduct is illegal.
Although the devil is in the details, in most cases the answer will likely be yes. What copyright law cares about, above all else, is preserving the incentive for human beings to create artistic and scientific works. Therefore, it is generally illegal to copy protected works without permission. And the doctrine of “fair use,” which provides a defense to certain claims of copyright infringement, typically doesn’t apply to copying that will significantly diminish the ability of copyright holders to make money from their works (thus significantly diminishing the incentive to create in the future). Generative AI has the potential to flood the market with endless amounts of images, songs, articles, books, and more. People can prompt generative AI models to produce these outputs using a tiny fraction of the time and creativity that would otherwise be required. So by training generative AI models with copyrighted works, companies are creating something that often will dramatically undermine the market for those works, and thus dramatically undermine the incentive for human beings to create things the old-fashioned way.
Take, for example, biographies. If a company uses copyrighted biographies to train a model, and if the model is thus capable of generating endless amounts of biographies, the market for many of the copied biographies could be severely harmed. Perhaps not the market for Robert Caro’s Master of the Senate, because that book is at the top of so many people’s lists of biographies to read. But you can bet that the market for lesser-known biographies of Lyndon B. Johnson will be affected. And this, in turn, will diminish the incentive to write biographies in the future.
Or take magazine articles. If a company uses copyrighted magazine articles to train a model capable of generating similar articles, it’s easy to imagine the market for the copied articles diminishing substantially. Especially if the AI-generated articles are made available for free. And again, how will this affect the incentive for human beings to put in the effort necessary to produce high-quality magazine articles?
With some types of works, the picture is a bit murkier. For example, it’s not clear how generative AI would affect the market for memoirs or autobiographies, since by definition people read those works because of who wrote them. With fiction, it might depend on the type of book. Perhaps classic works of literature like The Catcher in the Rye would not see their markets diminished. But the market for the typical human-created romance or spy novel could be diminished substantially by the proliferation of similar AI-created works. And again, the proliferation of such works would presumably diminish the incentive for human beings to write romance or spy novels in the first place.
Some students of copyright law respond that none of this matters because when companies use copyrighted works to train generative AI models, they are using the works in a way that’s highly creative in its own right. In the language of copyright law, the companies’ use of the works is “transformative.” As a factual matter, there’s no disputing that. And as a legal matter, it’s true that you’re less likely to be liable for copyright infringement if you’re copying the work for a transformative purpose. In that situation, you’re more likely to be protected by the fair use doctrine. But as the Supreme Court has emphasized, the fair use inquiry is highly fact dependent, and there are few bright-line rules. There is certainly no rule that when your use of a protected work is “transformative,” this automatically inoculates you from a claim of copyright infringement. And here, copying the protected works, however transformative, involves the creation of a product with the ability to severely harm the market for the works being copied, and thus severely undermine the incentive for human beings to create. Under the fair use doctrine, harm to the market for the copyrighted work is more important than the purpose for which the copies are made.
Speaking of which, in a recent ruling on this topic, Judge Alsup focused heavily on the transformative nature of generative AI while brushing aside concerns about the harm it can inflict on the market for the works it gets trained on. Such harm would be no different, he reasoned, than the harm caused by using the works for “training schoolchildren to write well,” which could “result in an explosion of competing works.” According to Judge Alsup, this “is not the kind of competitive or creative displacement that concerns the Copyright Act.” But when it comes to market effects, using books to teach children to write is not remotely like using books to create a product that a single individual could employ to generate countless competing works with a miniscule fraction of the time and creativity it would otherwise take. This inapt analogy is not a basis for blowing off the most important factor in the fair use analysis.
Another argument offered in support of the companies is more rhetorical than legal: Don’t rule against them, or you’ll stop the development of this groundbreaking technology. The technology is certainly groundbreaking. But the suggestion that adverse copyright rulings would stop this technology in its tracks is ridiculous. These products are expected to generate billions, even trillions, of dollars for the companies that are developing them. If using copyrighted works to train the models is as necessary as the companies say, they will figure out a way to compensate copyright holders for it.
The upshot is that in many circumstances it will be illegal to copy copyright-protected works to train generative AI models without permission. Which means that the companies, to avoid liability for copyright infringement, will generally need to pay copyright holders for the right to use their materials.
But that brings us to this particular case. …
The plaintiffs are thirteen published authors who have written, and who hold copyright in, various works. Those works are mostly novels, but also include plays, short stories, memoirs, essays, and nonfiction books. Examples include Sarah Silverman’s The Bedwetter, a comic memoir; Rachel Louise Snyder’s No Visible Bruises: What We Don’t Know About Domestic Violence Can Kill Us, a nonfiction book about domestic violence and how to combat it; Junot Díaz’s Pulitzer Prize–winning novel, The Brief Wondrous Life of Oscar Wao; and Andrew Sean Greer’s Less, also a Pulitzer Prize–winning novel. All of the books in which the plaintiffs hold copyright can be found in the datasets Meta downloaded, including both Books3 and the Anna’s Archive databases. In total, Meta downloaded at least 666 copies of books whose copyrights the plaintiffs hold.
The plaintiffs filed this lawsuit seeking to represent a class of all owners of copyrighted works used as training data for Llama.
… Given the state of the record, the Court has no choice but to grant summary judgment to Meta on the plaintiffs’ claim that the company violated copyright law by training its models with their books. But in the grand scheme of things, the consequences of this ruling are limited. This is not a class action, so the ruling only affects the rights of these thirteen authors—not the countless others whose works Meta used to train its models. And, as should now be clear, this ruling does not stand for the proposition that Meta’s use of copyrighted materials to train its language models is lawful. It stands only for the proposition that these plaintiffs made the wrong arguments and failed to develop a record in support of the right one.
I. COPYRIGHT LAW AND FAIR USE
Factor One
The first factor “considers the reasons for, and nature of, the copier’s use of an original work.” Primarily, the first factor focuses on whether the secondary use is “transformative”—that is, on whether and to what extent “the new work merely supersedes the objects of the original creation (supplanting the original), or instead adds something new, with a further purpose or different character.” Warhol, 598 U.S. at 528. Allowing a use with a “distinct purpose” is often consistent with the goals of copyright because it encourages the development of new expression “without diminishing the incentive to create.” Id. at 531. On the other hand, a secondary use with the same purpose as the original work is “more likely to provide the public with a substantial substitute for” the original. Id. at 531–32. This factor favors Meta. There is no serious question that Meta’s use of the plaintiffs’ books had a “further purpose” and “different character” than the books—that it was highly transformative. The purpose of Meta’s copying was to train its LLMs, which are innovative tools that can be used to generate diverse text and perform a wide range of functions. Cf. Oracle, 593 U.S. at 30 (transformative to use copyrighted computer code “to create a new platform that could be readily used by programmers”). Users can ask Llama to edit an email they have written, translate an excerpt from or into a foreign language, write a skit based on a hypothetical scenario, or do any number of other tasks. The purpose of the plaintiffs’ books, by contrast, is to be read for entertainment or education.
The plaintiffs do not meaningfully disagree about Llama’s purpose. To the contrary, they acknowledge that LLMs have “end uses” including serving “as a personal tutor,” assisting “with creative ideation,” and helping users “generate business reports.” And several of the plaintiffs testified to using LLMs for various purposes, all distinct from creating or reading an expressive work like a novel or biography—for instance, to find recipes, get tax or medical advice, translate documents, or conduct research. All of these functions are different from the use to which the plaintiffs’ books are generally put. So copying the books to develop a tool that can perform those functions is a use with a different purpose and character than the books themselves.
The plaintiffs’ law professor amici argue that Meta’s use has the same purpose and character as the books because an LLM training on a book is akin to a human reading one. One might also analogize Meta’s copying of the books to train Llama to a situation in which a professor copies a book and gives it to a student so that the student can use the knowledge from the book (along with knowledge they get from other books) to go do great things. But there are a few important differences.
First, an LLM’s consumption of a book is different than a person’s. An LLM ingests text to learn “statistical patterns” of how words are used together in different contexts. It does so by taking a piece of text from its training data, removing a word from that text, predicting what that word will be, and updating its general understanding of language based on whether it was right or wrong—and then repeating this exercise billions or trillions of times with different text. This is not how a human reads a book.
Second, unlike the hypothetical professor, Meta did not just give the plaintiffs’ books to one person. Meta copied the plaintiffs’ books as part of an effort to create a tool that can generate a wide range of text. Any person can use that tool to help them create further expression, whether by having it help them brainstorm or research for a creative writing project (like plaintiff David Henry Hwang, a playwright and screenwriter) or by having it write code to develop new software programs (like Lockheed Martin). By creating a tool that anyone can use, Meta’s copying has the potential to exponentially multiply creative expression in a way that teaching individual people does not. Cf. Oracle, 593 U.S. at 30.
In contrast to the copyright professors, the plaintiffs make different (and much weaker) arguments for why Meta’s use is not transformative. For example, the plaintiffs suggest that Llama has “no critical bearing” on their books, the way criticism or parody would. But “critique or commentary on the original” are not “the only uses that will furnish a justification ultimately qualifying as fair use.” Romanova, 138 F.4th at 115. To the contrary, a use that enables “the furnishing of valuable information on any subject of public interest” or renders “a valuable service to the public” might be justified, especially where that benefit is “provided without allowing public access to the copy.” Id.
In addition, the plaintiffs argue that Meta’s use is not transformative because Llama will output material that “mimics” the plaintiffs’ work or writing styles if prompted to do so. Therefore, the plaintiffs say, Meta’s use “merely amounts to a ‘repackaging’ “ of their books. The plaintiffs point to evidence that they say shows that Meta trained Llama to be able to emulate certain writers’ styles. But this evidence does not show that Meta trained Llama to repackage the plaintiffs’ works. To the contrary, as noted above, even using “adversarial” prompts designed to get Llama to regurgitate its training data, Llama will not produce more than 50 words of any of the plaintiffs’ books. And there is no indication that it will generate longer portions of text that would function as “repackaging” of those books. Nor is there even any indication that, as the plaintiffs’ amici claim, Meta developed Llama with the purpose of enabling it to create books that compete with the plaintiffs’ (without rising to the level of repackaging them). So at most, this evidence shows that Meta wanted Llama to be able to generate text in certain styles. But style is not copyrightable—only expression is. See 17 U.S.C. § 102(b); cf. Mattel, Inc. v. MGA Entertainment, Inc., 616 F.3d 904, 916 (9th Cir. 2010). Even if one possible use of Llama is to generate text with similarities to unprotectable aspects of the plaintiffs’ books, that does not mean Meta’s copying had the same purpose as those books.6
Footnote 6: By contrast, consider an LLM that was designed to be used to create works substantially similar to those on which it was trained, or to create works that competed with the originals without being substantially similar. Using copyrighted works to train such an LLM could be less transformative than using them to train a general-purpose LLM, because that use would have the purpose and character of enabling an LLM to develop substitute works. That said, even then, training the LLM would still likely be at least somewhat transformative; transformativeness isn’t an on-off switch.
FACTOR FOUR: THE EFFECT OF THE USE UPON THE POTENTIAL MARKET FOR OR VALUE OF THE COPYRIGHTED WORK
This factor looks to both the “extent of market harm caused by the particular actions of the alleged infringer” and to “whether unrestricted and widespread conduct of the sort engaged in by the defendant ... would result in a substantially adverse impact on the potential market’ for the original.” Campbell, 510 U.S. at 590). The “only harm” relevant to this factor “is the harm of market substitution.” Id. at 593. When, by contrast, the secondary work kills demand for the first through criticism or parody, the harm is not cognizable under the Copyright Act. Also relevant to this factor are “the public benefits the copying will likely produce.” Oracle, 593 U.S. at 35, 141 S.Ct. 1183.
As noted previously, the fourth factor is “undoubtedly the single most important element of fair use.” Harper & Row, 471 U.S. at 566. Meta is therefore wrong to suggest that, because the first factor strongly favors it, the inquiry should basically end there. To the contrary, given the fourth factor’s importance, it’s easy to imagine a situation in which a secondary use is highly transformative but the secondary user nonetheless loses on fair use because allowing people to engage in that kind of use would have too great an effect on the market for the original work. But by the same token, in a case where the first factor cuts strongly in favor of the defendant, generally the plaintiff’s only chance to defeat fair use will be to win decisively on factor four.
In a case involving the use of copyrighted works to train generative AI models, there are at least three ways a plaintiff might try to argue that the defendant’s copying harmed the market for the works (or that the market would be harmed if that copying were widespread). First, the plaintiff might claim that the model will regurgitate their works (or outputs that are substantially similar), thereby allowing users to access those works or substitutes for them for free via the model. Second, the plaintiff might point to the market for licensing their works for AI training and contend that unauthorized copying for training harms that market (or precludes the development of that market). Third, the plaintiff might argue that, even if the model can’t regurgitate their own works or generate substantially similar ones, it can generate works that are similar enough (in subject matter or genre) that they will compete with the originals and thereby indirectly substitute for them. In this case, the first two arguments fail. The third argument is far more promising, but the plaintiffs’ presentation is so weak that it does not move the needle, or even raise a dispute of fact sufficient to defeat summary judgment.
A
If Llama could be used to generate significant portions of the plaintiffs’ books—or text so similar to their books as to be infringing in its own right—that would threaten the market for the books because people would read those outputs instead. But that theory of harm is not viable in this particular case because, as discussed above, Llama does not allow users to generate any meaningful portion of the plaintiffs’ books. Neither party’s expert opined that Llama was able to regurgitate more than 50 words from any of the plaintiffs’ books, even in response to “adversarial” prompting designed specifically to make LLMs regurgitate. And the plaintiffs’ expert conceded that Llama would not generate “any significant percentage” of their books. In Google Books, by way of comparison, the Second Circuit held that the secondary use did “not threaten the rights holders with any significant harm to the value of their copyrights or diminish their harvest of copyright revenue” despite allowing users to see snippets adding up to as much as 16% of a book.11 804 F.3d at 224. Llama’s ability to regurgitate miniscule portions of the plaintiffs’ books if manipulated into doing so does not threaten to have a “meaningful or significant effect ‘upon the potential market for or value of’ “ the plaintiffs’ books. Id. (quoting 17 U.S.C. § 107(4)).
B
The plaintiffs’ primary theory of market harm is that Meta’s unauthorized use of their books for LLM training harms the market for licensing their books for that purpose. The plaintiffs devote nearly all of their discussion of the fourth factor to this theory. The parties therefore go back and forth at length about whether a market for licensing general trade books exists or is likely to develop.
But whether such a market exists or is likely to develop is irrelevant, because this market is not one that the plaintiffs are legally entitled to monopolize. In every fair use case, the “plaintiff suffers a loss of a potential market if that potential [market] is defined as the theoretical market for licensing” the use at issue in the case. Tresóna Multimedia, LLC v. Burbank High School Vocal Music Association, 953 F.3d 638, 652 (9th Cir. 2020)). Therefore, to prevent the fourth factor analysis from becoming circular and favoring the rightsholder in every case, harm from the loss of fees paid to license a work for a transformative purpose is not cognizable. Id.; Bill Graham Archives v. Dorling Kindersley Ltd., 448 F.3d 605, 614–15 (2d Cir. 2006); see also Oracle, 593 U.S. at 38, 141 S.Ct. 1183 (“cautioning against the ‘danger of circularity’”.
C
The third way that using copyrighted books to train an LLM might harm the market for those works is by helping to enable the rapid generation of countless works that compete with the originals, even if those works aren’t themselves infringing. Assume for this discussion that people can (or will soon be able to) use LLMs to generate massive amounts of text in significantly less time than it would take to write that text, and using a fraction of the creativity. People could thus use LLMs to create books and then sell them, competing with books written by human authors for sales and attention. Indeed, to some extent, this appears to be occurring already—one expert for the plaintiffs briefly discusses reports of AI-generated books “flooding Amazon.” People might even be motivated to make those books available for free, given how easily it will presumably be to prompt an LLM to create them. Harm from this form of competition is the harm of market dilution. Or as one commentator describes it, the harm of “indirect” substitution, rather than “direct” substitution (which would be the first form of harm described). See Matthew Sag, Fairness and Fair Use in Generative AI, 92 Fordham Law Review 1887, 1916–20 (2024).
Of course, not all copyrighted works would have their markets diluted equally by AI-generated competitors. It seems unlikely, for instance, that AI-generated books would meaningfully siphon sales away from well-known authors who sell books to people looking for books by those particular authors. But it’s easy to imagine that AI-generated books could successfully crowd out lesser-known works or works by up-and-coming authors. While AI-generated books probably wouldn’t have much of an effect on the market for the works of Agatha Christie, they could very well prevent the next Agatha Christie from getting noticed or selling enough books to keep writing.12
Footnote 12: To be clear, the point is not that authors are entitled to more or less copyright protection based on how famous or popular they are. Cf. Warhol, 598 U.S. at 544 & n.19. The point is that different works may have different markets that will be affected differently by floods of AI-generated competitors.
This effect also seems likely to be more pronounced with respect to certain types of works. For instance, an AI model that can generate high-quality images at will might be expected to greatly affect the market for such images, diminishing the incentive for humans to create them. An LLM that could generate accurate information about current events might be expected to greatly harm the print news market. The market for certain nonfiction works—for example, books about how to take care of your garden—could be greatly diminished by the ability of LLMs to produce books on that topic. For fiction works, it might be more dependent on the author or the genre in which that author operates.
The difference might be in part because some works are relatively functional and generally less dependent on the author’s creativity. When picking a news article, readers want something that will tell them about a current (or past) event clearly, accurately, and concisely. When picking a novel, by contrast, readers may care about a much longer list of characteristics. They may care, for instance, about tone, thematic depth, writing style, plot, or characters; they may want a book that contains a number of plot twists or depicts a certain type of character development. These elements of a novel depend greatly on the creativity of the author. While a news article is also a product of its author’s creativity (especially with respect to things like structure and diction), there are many more creative choices in the average novel than the average news article, and those creative choices are more important to the average novel’s quality. Relatedly, one could imagine people caring more about whether a novel is AI-generated (as opposed to the product of human creativity) than whether a news article is AI-generated.13
Footnote 13: This is not to suggest that news articles or other works that may be less dependent on their author’s creativity are thus less deserving of protection, or that it would therefore be more appropriate to use those works to train an LLM. To the contrary, as noted with respect to the second factor, nonfiction works are still protected by copyright because the law protects their authors’ choices as to how to express facts. See Google Books, 804 F.3d at 220.
It also should be noted that, when considering market dilution, the proper comparison isn’t to a world with no LLMs, but to a world where LLMs weren’t trained on copyrighted books. Perhaps an LLM trained only on public domain works could still be capable of quickly generating large numbers of books that could compete for sales with copyrighted books. But there is plenty of evidence in the record that training on books substantially benefits LLMs’ creativity and ability to generate long pieces of text. And because LLMs perform better the more text they are trained on, an LLM trained only on public domain books would presumably, all else equal, lag significantly behind a book trained also on copyrighted ones. So training an LLM on copyrighted books would seem, in most circumstances, to make that LLM better able to generate works that could dilute the market for the books in its training data.
Meta and its law professor amici, as well as the Matthew Sag article cited above, argue that market dilution does not count under the fourth factor. They argue that harm caused by an LLM’s outputs is only relevant if the outputs are themselves infringing—that is, if the LLM regurgitates copyrighted material (or generates text that is substantially similar to copyrighted material). But that can’t be right. To be sure, it would be easier to conclude that the market for copied books would be harmed by an LLM that is capable of regurgitating those books or generating substantially similar text. But less similar outputs, such as books on the same topics or in the same genres, can still compete for sales with the books in the training data. And by taking sales from those books, or by flooding stores and online marketplaces so that some of those books don’t get noticed and purchased, those outputs would reduce the incentive for authors to create—the harm that copyright aims to prevent.
The Supreme Court has said that the “only harm” that matters under the fourth factor “is the harm of market substitution.” Campbell, 510 U.S. at 593. But indirect substitution is still substitution: If someone bought a romance novel written by an LLM instead of a romance novel written by a human author, the LLM-generated novel is substituting for the human-written one. This is different from the (non-cognizable) harm caused by criticism or commentary, which can harm demand for an original work without serving as a replacement for it.
Relatedly, Meta argues that “legitimate” competition from noninfringing secondary works is not cognizable under the fourth factor. It cites the intermediate copying cases for this proposition. See Sega, 977 F.2d at 1523–24; Connectix, 203 F.3d at 607. But key to those cases’ reasoning was the fact that the secondary users’ competing products did not benefit from the creative expression in the works they copied. By contrast, as discussed, LLMs are better able to generate text (including competing works) because they are trained on the creative expression in copyrighted books. So this competition is not “legitimate” within the meaning of those cases.
It’s true that, in many copyright cases, this concept of market dilution or indirect substitution is not particularly important. That’s because, in a more typical case, an original work is being compared to a single secondary work. If the secondary work is somewhat similar, but not so similar as to effectively be a copy, it still might have a small indirect effect on the market for the original work. But that likely won’t matter. Recall that the fourth factor looks to whether “conduct of the sort engaged in by the defendant” would have a “substantially adverse impact on the potential market for the original.” Campbell, 510 U.S. at 590. The existence of some harm from indirect substitution isn’t dispositive of the fourth factor or the fair use inquiry. Where, for instance, the first factor cuts in favor of the secondary user, the law might tolerate a little bit of competition. See Google Books, 804 F.3d at 224. In cases involving a single secondary work that’s similar-but-not-too-similar, it’s unlikely that harm from market dilution would be significant enough to matter. Even considering the effect of “widespread conduct of the sort engaged in by the defendant,” Oracle, 593 U.S. at 38, creating one indirectly substitutional work at a time could only have so great an effect on the market for the original.
This case is different. This is not a case where an original work is being compared to one secondary work. Nor is this case like the previous fair use cases involving creation of a digital tool. In those cases, like Google Books and Perfect 10, the tool could at most be used to access part or all of the original works. This case, unlike any of those cases, involves a technology that can generate literally millions of secondary works, with a miniscule fraction of the time and creativity used to create the original works it was trained on. No other use—whether it’s the creation of a single secondary work or the creation of other digital tools—has anything near the potential to flood the market with competing works the way that LLM training does. And so the concept of market dilution becomes highly relevant.
In arguing that this sort of harm doesn’t count just because it’s never made a difference in a case before, Meta makes the mistake the Supreme Court instructs parties and courts to avoid: robotically applying concepts from previous cases without stepping back to consider context. Fair use is meant to be a flexible doctrine that takes account of “significant changes in technology.” Oracle, 593 U.S. at 19, 141 S.Ct. 1183 (quoting Sony, 464 U.S. at 430). Courts can’t stick their heads in the sand to an obvious way that a new technology might severely harm the incentive to create, just because the issue has not come up before. Indeed, it seems likely that market dilution will often cause plaintiffs to decisively win the fourth factor—and thus win the fair use question overall—in cases like this.
But courts can’t decide cases based on what they think will or should happen in other cases. They must decide cases based on the arguments presented and the evidence submitted by the parties. The question, then, is whether these particular thirteen plaintiffs in this particular case have presented enough evidence to win on this factor. Or, to put it more precisely given the procedural posture of this case, whether these plaintiffs have presented enough evidence to raise a genuine dispute of material fact sufficient to give the question of market dilution to a jury. The answer is no.
In their complaint, the plaintiffs asserted only two types of market harm—that users of Llama can reproduce text from their books, and that Meta’s copying harmed the market for licensing copyrighted materials to companies for AI training. As for market dilution—the notion that allowing companies like Meta to copy their works to train products like Llama would inevitably cause the market for the plaintiffs’ works to be flooded with similar works—the plaintiffs never so much as mentioned it in their complaint. Nor did they mention it in their own summary judgment motion.
Naturally, given the allegations in the complaint, Meta’s cross-motion for summary judgment focused on defeating the first two theories. But Meta also noted in its motion that the plaintiffs hadn’t presented any evidence that Meta’s use of their books to train Llama had harmed book sales. And Meta presented its own expert testimony explaining that Llama 3’s release did not have any discernible effect on the plaintiffs’ sales (or those of other books in Llama’s training data), at least in the period shortly after the release.
In opposition, the plaintiffs’ primary response was that this was beside the point because of their first two theories. They did make fleeting reference to a report by one of their experts, who briefly discussed the concept of indirect substitution and mentioned articles discussing how AI-created books are starting to flood Amazon. But this discussion generates more questions than answers.
First, is Llama capable of generating such books? If it isn’t currently, will it be capable of doing so in the near future? Presumably the answer is yes, but that’s not a foregone conclusion. An LLM could, for instance, be configured to be unable to produce book-length or book-style outputs. So the fact that books are being created by some LLM does not automatically mean that Llama can create them or will be able to do so soon.
Second, what are these AI-generated books? Do they compete with Sarah Silverman’s memoir? With plaintiff Matthew Klam’s book of short stories? With Rachel Louise Snyder’s nonfiction work on domestic violence? The plaintiffs provide no analysis of the markets for their books, no discussion of whether these markets are or could be affected by AI-generated books, and no explanation of whether the existing AI-generated books referenced in the expert report compete in these markets.
Third, what impact does this competition actually have on sales of the books it competes with? Does it drown out those books entirely? Does it just chisel at their sales at the margins? Or, as discussed above and seems likely, does it depend on the book—are readers of romance novels happy to buy AI-generated ones, while all the people who want to read Sarah Silverman’s memoir still want to read it over AI-generated comic memoirs? Whatever the effects have been thus far, are they likely to increase in the future, as more and more AI-generated books are written, and as LLMs get better and better at writing human-like text?
Fourth, how does the threat to the market for the plaintiffs’ books in a world where LLM developers can copy those books compare to the threat to the market for the plaintiffs’ books in a world where the developers can’t copy them? There is no hint of that in the briefs or evidence presented by the plaintiffs.
The analysis is complicated somewhat by the fact that fair use is an affirmative defense and that Meta moved for summary judgment on it. For those reasons, Meta had the burden of presenting evidence that its copying doesn’t threaten to substantially harm the market for the plaintiffs’ books. It didn’t conclusively establish that its copying couldn’t do so in the future—potentially because its copying did in fact make Llama better able to generate countless works that will dilute the market for the plaintiffs’ books. But where a defendant introduces evidence of a lack of market harm, “and the plaintiff fails to introduce empirical evidence countering such a showing, the fourth factor should be weighed in the defendant’s favor.” Patry on Fair Use § 6:13. That is exactly what happened here. Meta introduced evidence that its copying hasn’t caused market harm. The plaintiffs presented no empirical evidence to the contrary—no evidence that the copying has already caused market harm, and no evidence that the copying is likely to cause market harm in the future. All the plaintiffs presented is speculation, and speculation is insufficient to raise a genuine issue of fact and defeat summary judgment.
The plaintiffs argue that they didn’t need to present empirical evidence because market harm can be inferred. For this argument, they cite to Hachette, in which the Second Circuit inferred market harm—even though the plaintiffs had not provided “empirical data” showing any and the secondary user presented expert testimony that there was none—because it was “self-evident” that the secondary use would cause such harm if widespread. 115 F.4th at 192–93. In Hachette, the secondary user maintained a database that let internet users “download an identical copy of” the plaintiffs’ books for free. Id. at 194. The secondary use therefore offered a directly “competing substitute” for the original books. Id. at 195.
While it made sense to infer market harm in Hachette, it doesn’t make sense to do so here. First, the Supreme Court has stated that no “inference of market harm ... is applicable to a case involving something beyond mere duplication for commercial purposes.” Campbell, 510 U.S. at 591. In Hachette, the secondary use was basically “mere duplication.” Here, by contrast, Meta’s use is highly transformative and has a purpose well beyond that. Second, unlike in Hachette, Meta’s use does not let users access any significant portion of the plaintiffs’ books, so it isn’t self-evident that Meta’s use would create harm via direct substitution. Nor is it self-evident that Llama will harm the book sale market by enabling users to create a flood of competing books. It’s possible, even likely, that Llama will harm the book sale market. But to conclude that it will requires inferring that Llama (and not just any LLM) can be used to create such books, that it will be used to create such books, that consumers will purchase those books instead of books written by human authors, that consumers will buy those books instead of the plaintiffs’ books in particular, and that Llama is meaningfully better at creating those books because it was trained on copyrighted material. In Hachette, on the other hand, the only necessary inference was that readers might choose to download the plaintiffs’ books for free instead of paying for them—a much shorter (and more obvious) inferential leap.
On this record, then, Meta has defeated the plaintiffs’ half-hearted argument that its copying causes or threatens significant market harm. That conclusion may be in significant tension with reality, but it’s dictated by the choice the plaintiffs made to put forward two flawed theories of market harm while failing to present meaningful evidence on the effect of training LLMs like Llama with their books on the market for those books.14
Footnote 14: The plaintiffs also assert that the market for their works was harmed in the more narrow sense that, if Meta had not downloaded the books from a shadow library, it would have been required to buy the books. But as already discussed, even though that downloading is a separate use, it must be considered in light of its overall purpose. For instance, imagine a researcher who downloaded books from a shadow library in the process of writing an article on shadow libraries, and only did so for their research. That downloading would almost certainly be a fair use. Of course, in that example, the downloader has less ability to procure the books elsewhere than Meta did. But the point is that downloading from a shadow library, which the plaintiffs refer to as “unmitigated piracy,” must be viewed in light of its ultimate end. Because Meta’s purpose of LLM training is so transformative, the plaintiffs needed to win decisively on the fourth factor. The loss of isolated sales to AI developers is not the kind of market harm that could tip the scales for the plaintiffs.
VII. CONCLUSION
Fair use is a fact-specific doctrine that requires case-by-case analysis that is sensitive to new technologies and their potential consequences. No previous case has involved a use that is both as transformative and as capable of diluting the market for the original works as LLM training is. So no previous case answers the question whether Meta’s copying was fair use. That question must be answered by flexibly applying the fair use factors and considering Meta’s copying in light of the purpose of copyright and fair use: protecting the incentive to create by preventing copiers from creating works that substitute for the originals in the marketplace.
In cases involving uses like Meta’s, it seems like the plaintiffs will often win, at least where those cases have better-developed records on the market effects of the defendant’s use. No matter how transformative LLM training may be, it’s hard to imagine that it can be fair use to use copyrighted books to develop a tool to make billions or trillions of dollars while enabling the creation of a potentially endless stream of competing works that could significantly harm the market for those books. And some cases might present even stronger arguments against fair use. For instance, as discussed above, it seems that markets for certain types of works (like news articles) might be even more vulnerable to indirect competition from AI outputs. On the other hand, though, tweak some facts and defendants might win. For example, using copyrighted books to train an LLM for nonprofit purposes, like national security or medical research, might be fair use even in the face of some amount of market dilution. See Oracle, 593 U.S. at 32 (“[A] finding that copying was not commercial in nature tips the scales in favor of fair use.”). Or plaintiffs whose works are unlikely to face meaningful competition from AI-generated ones may be unable to defeat a fair use defense.
In this case, because Meta’s use of the works of these thirteen authors is highly transformative, the plaintiffs needed to win decisively on the fourth factor to win on fair use. See, e.g., Perfect 10, 508 F.3d at 1168 (fair use where secondary use was “significant[ly] transformative” and fourth factor “favor[ed] neither party”). And to stave off summary judgment, they needed to create a genuine issue of material fact as to that factor. Because the issue of market dilution is so important in this context, had the plaintiffs presented any evidence that a jury could use to find in their favor on the issue, factor four would have needed to go to a jury. Or perhaps the plaintiffs could even have made a strong enough showing to win on the fair use issue at summary judgment. But the plaintiffs presented no meaningful evidence on market dilution at all. Absent such evidence and in light of Meta’s evidence, the fourth factor can only favor Meta. Therefore, on this record, Meta is entitled to summary judgment on its fair use defense to the claim that copying these plaintiffs’ books for use as LLM training data was infringement.
IT IS SO ORDERED.
Notes and questions
(1) The only authority Judge Chhabria cites for treating non-infringing competition (“market dilution”) as a cognizable harm under the fourth fair use factor is Matthew Sag, Fairness and Fair Use, but although that article articulates a framework for that kind of analysis, it also notes several reasons why courts might not want to go in that direction:
However, the difficulty with such an argument is that it tends to blur the line between copyright and unfair competition. Copyright law not only allows competition based on the ideas and unprotectable elements in a work; copyright law encourages it. In general, to say that an aspect of a work is uncopyrightable is to say that it should be the subject of free competition.
(2) The USCO made a similar market dilution argument in Copyright and Artificial Intelligence, Part 3: Generative AI Training, the pre-publication version of 9 May 2025, which is still the only version the Office has released: “Even where a model’s outputs are not substantially similar to any specific copyrighted work, they can dilute the market for works similar to those found in its training data, including by generating material stylistically similar to those works.” P. 73. Why didn’t Judge Chhabria cite the Copyright Office report?
(3) Despite his misgivings, Judge Chhabria ruled in favor of Meta on summary judgement after chastening the plaintiffs’ lawyers for not providing evidence of market dilution. Is it right to blame the lawyers for failing to find concrete and particularized evidence of market dilution, or should this have been a hint that such fears are more conjecture than reality?
(4) How is the harm Judge Chhabria predicts different to or similar to harms that courts have recognized in past fair use cases? Do you agree with Judge Alsup in the Anthropic case, who said that complaining of competition from works that did not contain an author’s original expression was “no different” to complaining “that training schoolchildren to write well would result in an explosion of competing works. This is not the kind of competitive or creative displacement that concerns the Copyright Act. The Act seeks to advance original works of authorship, not to protect authors against competition.” Or do you think that Judge Chhabria is right to worry about AI competition as a threat to authorship?
(5) The nonexpressive use framing is not universally accepted, and now is the time to weigh the objections against the cases you have read. David Opderbeck argues that the holdings of the nonexpressive use cases do not generalize in the way this textbook suggests: “It is not so clear, however, whether existing doctrine says anything so broad about bulk non-expressive uses.” Do Bartz and Kadrey bear that out? Both judges reached for the nonexpressive use idea without much hesitation, but neither had to decide whether it generalizes beyond training.
(6) Responding to Sag’s Copyright and Copy-Reliant Technology, James Grimmelmann argues in Copyright for Literate Robots that broadly applying a principle of nonexpressive use would privilege robotic readers over human ones. He concedes the value of digital humanities research but cautions that “not all robotic reading is so benign, and the logic of nonexpressive use encourages the circulation of copyrighted works in an underground robotic economy.” The shadow library material in the next section is that worry made concrete. Does Judge Alsup’s treatment of Anthropic’s pirated library answer Grimmelmann, or confirm him?
(7) In Copyright and the Training of Human Authors and Generative Machines, Robert Brauneis argues that robot readers should not be able to learn for free when human authors pay, directly or indirectly, for access to copyrighted works. That is close to the argument Judge Chhabria found attractive and Judge Alsup rejected. Which of them has the better of it, and does your answer depend on whether the books were bought or pirated?
(8) Benjamin Sobel, in Artificial Intelligence’s Fair Use Crisis, questions whether the nonexpressive use rationale should apply to “digital artifacts that are broadly equivalent to copyrightable human expression.” Sag answers in Copyright Safety for Generative AI that a model’s creation of pseudo-expression does not affect its nonexpressive use status: a use is nonexpressive when it is an intermediary step in a technical process that does not convey pre-existing expression to a new public. Notice that both Bartz and Kadrey proceed on the assumption that no output infringed. Sobel’s objection is about the cases where that assumption fails. How much of the reasoning in these two decisions survives if it does?
Lawful access and shadow libraries
In addition to the issues discussed above, the Bartz and Kadrey cases also diverged on the issue of “shadow libraries.” It is widely assumed that all the main LLM developers have at some point trained their models on works downloaded from sites of known infringement, or so-called shadow libraries, like Library Genesis and Sci-Hub. As you will see in the extract below, for Judge Alsup in Bartz v. Anthropic, downloading source copies from pirate sites is a wrong that can’t be righted by a subsequent transformative use. In Kadrey Judge Chhabria, in contrast, looked at Meta’s decision to collect training data from pirate sites as simply one step in a transformative process that he ultimately, if somewhat grudgingly, held was fair use:
Because Meta’s ultimate use of the plaintiffs’ books was transformative, so too was Meta’s downloading of those books. … [Even the unused] downloads the plaintiffs identify had the ultimate purpose of LLM training.
Judge Chhabria noted however that
downloading copyrighted material from shadow libraries would be relevant if it benefitted those who created the libraries and thus supported and perpetuated their unauthorized copying and distribution of copyrighted works. … But the plaintiffs have not submitted any evidence about this.
The division between the two judges is now teed up for the Ninth Circuit, but it has not yet been certified for review. Judge Chhabria denied the plaintiffs’ request to certify an interlocutory appeal on 8 July 2026, so the question will not reach the Court of Appeals until final judgment. In the meantime the plaintiffs’ remaining route against Meta is not the downloading itself but the seeding that accompanied it: the allegation that Meta, in torrenting the books, simultaneously distributed them to other users of the network. That theory does not depend on characterizing the download as an input to a transformative use, which is what defeated the plaintiffs on the copying claim, and it is worth asking why it should make a difference.
Bartz v. Anthropic PBC, 2025 WL 1741691 (N.D. Cal. June 23, 2025) (second extract)
This extract should be read in conjunction with the previous extract of this case, focusing on whether LLM training is fair use. Even though Anthropic won on that issue, it did not fare so well on another issue.
[Judge Alsup went into great detail on the how Anthropic came to the decision to avail itself of pirated books and how the company used the books it collected. The highlights of that discussion were that Anthropic deliberately chose piracy over legitimate purchasing to avoid “legal/practice/business slog,” systematically downloading over seven million pirated books from known illegal sources including Books3, LibGen, and PiLiMi between 2021-2022. Despite becoming concerned about legal risks, the company retained these pirated works and later spent millions purchasing physical books only to destructively scan them into digital copies, creating a massive “research library” designed to store everything “forever” for training AI models and other research purposes, with the ultimate goal of improving Claude’s performance for paying customers. Judge Alsup also highlighted the contradiction in Anthropic’s defense strategy—while arguing that pirating millions of books was justified as “reasonably necessary” for AI training, the company has simultaneously resisted disclosure and even clawed back evidence showing which specific copies were actually used for training, forcing the court to hold these evidentiary deficiencies against Anthropic rather than the plaintiffs.]
ANALYSIS
Section 107 of the Copyright Act identifies four factors for determining whether a given use of a copyrighted work is a fair use … These factors presuppose a “use.” So, at the threshold, a court must decide whether a “copyrighted [work] has been used in multiple ways,” then evaluate each. Warhol, 598 U.S. at 533. Uses do not turn on “the subjective intent of the user” but on “an objective inquiry into what use was made, i.e., what the user d[id] with the original work.” Id. at 544–45. A “use” should be construed narrowly enough to not “swallow” distinguishable infringing uses, much less categories of exclusive rights in toto. Sometimes, the challenged copying involves just one use: In Perfect 10, Inc. v. Amazon.com, Inc., Google visited websites having full-sized images, made only reduced-sized copies, and incorporated those directly into its search engine — the sole use of the thumbnails being as “pointer[s]” to the images themselves. 508 F.3d 1146, 1157, 1160, 1165 (9th Cir. 2007). Sometimes, the copying involves many uses: In the Google Books cases, Google borrowed books from libraries, made both full-image and text-only copies, and incorporated different copies into different tools — one use being to reveal information “about those books,” another use being to provide the books to print-disabled patrons, and still another being to back up the print books if lost. [Citing Google Books and HathiTrust]
Our parties debate an instructive decision. In American Geophysical Union v. Texaco Inc., Texaco employees used scientific articles in a central library, used copies of them in personal desk libraries, and used selected copies again in the scientific laboratory — the first use paid for, the second infringing, and the third plausibly fair but in fact a rare occurrence.
Here, our parties contest what use or uses are at issue. Anthropic contends it copied Authors’ books only for one use: Only to train LLMs. By contrast, Authors contend it did so for at least two uses: First to build a vast, central library of potentially useful content, and second to train specific LLMs using shifting sets and subsets of that content — over time selecting the more well-organized and well-expressed works for training. Authors also complain that the print-to-digital format change was itself an infringement not abridged as a fair use. Authors do not allege, however, that any LLM outputs infringing upon their works ever reached users of the public-facing Claude service.
This order addresses each of the four factors in turn, pointing out how each applies to the training copies and to the purchased and pirated library copies. It concludes with an integrated analysis.
1. The Purpose and Character of the Use.
… B. The Copies Used to Build a Central Library.
Recall that Anthropic purchased millions of print books for its central library and pirated millions of digital books for its central library, too. It used specific sets and subsets of books for training specific LLMs. And, it then retained all the copies in its central library for other uses that might arise even after deciding it would not use them to train any LLM (at all or ever again). Anthropic seems to believe that because some of the works it copied were sometimes used in training LLMs, Anthropic was entitled to take for free all the works in the world and keep them forever with no further accounting. There is no carveout, however, from the Copyright Act for AI companies.
Because the legal issues differ between the library copies Anthropic purchased and pirated, this order takes them in turn.
(i) The Purchased Library Copies Converted from Print to Digital.
Anthropic purchased millions of print copies to “build a research library”. It destroyed each print copy while replacing it with a digital copy for use in its library (not for sharing nor sale outside the company). As to these copies, Authors do not complain that Anthropic failed to pay to acquire a library copy. Authors only complain that Anthropic changed each copy’s format from print to digital. On the facts here, that format change itself added no new copies, eased storage and enabled searchability, and was not done for purposes trenching upon the copyright owner’s rightful interests — it was transformative.
Anthropic purchased its print copies fair and square. With each purchase came entitlement for Anthropic to “dispose[ ]” each copy as it saw fit. 17 U.S.C. § 109(a). So, Anthropic was entitled to keep the copies in its central library for all the ordinary uses. Yes, Anthropic changed the format of these library copies from print to digital — giving rise to the issue here.
All agree on the facts of the format change. Anthropic “destructively scanned” the print copies to create the digital ones. Anthropic or its vendors stripped the bindings from the print books, cut the pages to workable dimensions, and scanned those pages — discarding each print copy while creating a digital one in its place. The digital copy was then housed in the “research library” or “generalized data area” in place of the print copy. Authors do not allege and our record does not show that Anthropic provided its converted digital copies of print books to anyone outside Anthropic.
The parties disagree about the legal consequences of the format change. Was scanning the print copies to create digital replacements transformative? Anthropic argues it was because it was reasonably necessary to training LLMs. Authors argue it was a distinguishable step requiring independent justification.
Here, for reasons narrower than Anthropic offers, the mere format change was a fair use.
Storage and searchability are not creative properties of the copyrighted work itself but physical properties of the frame around the work or informational properties about the work. See Texaco, 802 F. Supp. at 14 (physical), aff’d, 60 F.3d at 919; Google, 804 F.3d at 225 (informational); Sony Corp. of Am. v. Universal City Studios, Inc. (“Sony Betamax”), 464 U.S. 417, 447 (1984) (rightful interests). In Texaco, the court reasoned that if a purchased scientific journal article had been copied “onto microfilm to conserve space, this might [have been] a persuasive transformative use.” 802 F. Supp. at 14, aff’d, 60 F.3d at 919 (reducing “bulk” “might suffice to tilt the first fair use factor in favor of Texaco if these purposes were dominant”). In Google Books, the court reasoned that a print-to-digital change to expose information about the work was transformative. Google, 804 F.3d at 225. And, in Sony Betamax, the Supreme Court held that making a recording of a television show in order to instead watch it at a later time was copying but did not usurp any rightful interest of the copyright owner. 464 U.S. at 447. Important to the Supreme Court’s reasoning was the expectation that most such copiers would not distribute the permanent copies of the work. Finally, in A&M Records, Inc. v. Napster, Inc., our court of appeals recognized the reasoning just explained, and therefore rejected by contrast a digitization effort that was touted as space-shifting but in fact resulted in the multiplication of copies shared with outsiders through a file-sharing service.
Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others).
Yes, Anthropic is a commercial outfit. And, this order takes for granted that Anthropic in fact benefited from the print-to-digital format change — or it would not have gone to all the trouble. But the crux of the first fair use factor’s concern for “commercial” use is in protecting the copyright owners and their entitlements to exploit their copyright as they see fit (or not). See, e.g., Harper & Row, Publishers, Inc. v. Nation Enters., 471 U.S. 539, 562 (1985). That the accused is a commercial entity is indicative, not dispositive. That the accused stands to benefit is likewise indicative. But what matters most is whether the format change exploits anything the Copyright Act reserves to the copyright owner. Anthropic already had purchased permanent library copies (print ones). It did not create new copies to share or sell outside.
Yes, Authors also might have wished to charge Anthropic more for digital than for print copies. And, this order takes for granted that Authors could have succeeded if Anthropic had been barred from the format change. “But the Constitution’s language [in Clause 8] nowhere suggests that [the copyright owner’s] limited exclusive right should include a right to divide markets or a concomitant right to charge different purchasers different prices for the same book, [merely] say to increase or to maximize gain.” See Kirtsaeng v. John Wiley & Sons, Inc., 568 U.S. 519, 552 (2013). Nor does the Copyright Act itself. Section 106 sets out exclusive rights that fair uses under Section 107 abridge. Section 106(1) reserves to the copyright owner the right to make reproductions. But on our facts we face the unusual situation where one copy entirely replaced the another. And, Section 106(2) reserves to the copyright owner the right to make derivative works that add or subtract creative material — as occurs in a “translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, [or] condensation” of a book, 17 U.S.C. § 101 (definitions). For some “other modification[ ]” of a book to constitute a “derivative work,” it must itself “represent an original work of authorship.” Ibid. But on our facts the format was changed but no content was added or subtracted. Section 106(3) further reserves to the copyright owner the right to distribute copies. But again, the replacement copy here was kept in the central library, not distributed.
As a result, Anthropic’s format-change from print library copies to digital library copies was transformative under fair use factor one. Anthropic was entitled to retain a copy of these works in a print format. It retained them instead in a digital format, easing storage and searchability. And, the further copies made therefrom for purposes of training LLMs were themselves transformative for that further reason, as above.
To be clear, this print-to-digital conversion involved a different and narrower form of transformative use than the broader one advanced by Anthropic. Anthropic argues that the central library use was part and parcel of the LLM training use and therefore transformative. This order disagrees. However, this order holds that the mere conversion of a print book to a digital file to save space and enable searchability was transformative for that reason alone. Therefore, the digital copy should be treated just as if the purchased print copy had been placed in the central library.
In sum, the first fair use factor favors fair use for the digital library copies converted from purchased print library copies — but these do not excuse the pirated library copies.
(ii) The Pirated Library Copies.
Before buying books for its central library, Anthropic downloaded over seven million pirated copies of books, paid nothing, and kept these pirated copies in its library even after deciding it would not use them to train its AI (at all or ever again). Authors argue Anthropic should have paid for these pirated library copies. This order agrees.
The basic problem here was well-stated by Anthropic at oral argument: “You can’t just bless yourself by saying I have a research purpose and, therefore, go and take any textbook you want. That would destroy the academic publishing market if that were the case”. Of course, the person who purchases the textbook owes no further accounting for keeping the copy. But the person who copies the textbook from a pirate site has infringed already, full stop. This order further rejects Anthropic’s assumption that the use of the copies for a central library can be excused as fair use merely because some will eventually be used to train LLMs.
This order doubts that any accused infringer could ever meet its burden of explaining why downloading source copies from pirate sites that it could have purchased or otherwise accessed lawfully was itself reasonably necessary to any subsequent fair use. There is no decision holding or requiring that pirating a book that could have been bought at a bookstore was reasonably necessary to writing a book review, conducting research on facts in the book, or creating an LLM. Such piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded.
But this order need not decide this case on that rule. Anthropic did not use these copies only for training its LLM. Indeed, it retained pirated copies even after deciding it would not use them or copies from them for training its LLMs ever again. They were acquired and retained, as a central library of all the books in the world.
Building a central library of works to be available for any number of further uses was itself the use for which Anthropic acquired these copies. One further use was making further copies for training LLMs. But not every book Anthropic pirated was used to train LLMs. And, every pirated library copy was retained even if it was determined it would not be so used. Pirating copies to build a research library without paying for it, and to retain copies should they prove useful for one thing or another, was its own use — and not a transformative one.
Anthropic’s briefing contains other reasons why it believes its pirated library copies are irrelevant to our fair use analysis, notwithstanding its own statements at our oral argument.
First, Anthropic accepts in this posture that it acted in bad faith but argues that its bad faith in pirating copies cannot “somehow short-circuit” the fair use analysis (Reply 6 (downplaying Atari Games Corp. v. Nintendo of Am., Inc., 975 F.2d 832, 843 (Fed. Cir. 1992) (applying law of Ninth Circuit))). But its bad faith is not the basis for this decision. Each use of a work must be analyzed objectively. Warhol, 598 U.S. at 544–45. The objective analysis here shows the initial copies were pirated to create a central, general-purpose library, as a substitute for paid copies to do the same thing. Of course, if infringement is found, bad faith would matter for determining willfulness.
Second, Anthropic argues that its goal to put the copies eventually “to a highly transformative use” requires that each copy and use along the way be justified as having a transformative use, too. But now Anthropic seeks to take the shortcut Anthropic just said cannot be taken. Again, the Supreme Court tasks us with looking past the “subjective intent of the user” to the objective use made of each copy. Warhol, 598 U.S. at 544–45 (emphasis added). Put another way, what a copyist says or thinks or feels matters only to the extent it shows what a copyist in fact does with the work. Indeed, the same copy can be used one way, then another, each with a different result. Id. at 533. Here, what Anthropic said about its acquisitions at the time — that they were made to “build[ ] a research library” while avoiding a “huge legal/practice/business slog” — are relevant in this regard. And, Anthropic’s actual use of these pirated copies was to create its central library of texts that, like any university or corporate library, stored the works’ well-organized facts, analyses, and expressive examples for various contingent uses, one being training.5
Footnote 5: Our court of appeals has not yet reappraised how bad faith (or good faith) figures in fair use after Warhol. Its prior appraisal applied the Supreme Court’s statement that “[f]air use presupposes good faith and fair dealing,” Harper & Row, 471 U.S. at 562, 105 S.Ct. 2218 (cleaned up). See Perfect 10, 508 F.3d at 1164 n.8. Since then, the Supreme Court has renewed its “skepticism about whether bad faith has any role.” Oracle, 593 U.S. at 32–33, 141 S.Ct. 1183 (reiterating doubts of Campbell, 510 U.S. at 585 n.18). And, recently, the Supreme Court has held squarely that it is not the “subjective intent” of a copyist that counts, but the “objective ... use” of the copy. Warhol, 598 U.S. at 544–45.
Third, Anthropic argues that Texaco — the case involving copies used in a central library, copies used in desk libraries, and copies used in the laboratory — is inapposite. Anthropic argues that the disputed copies in Texaco were never used in the laboratory but instead in personal desk libraries for a use “identical to the original purpose and use” of the central library copies, and so not for a transformative use. By contrast, says Anthropic, here it did use copies in the laboratory to train LLMs — a very transformative use. But this is a fast glide over thin ice. Like Texaco, Anthropic possessed copies it did not put into use in the laboratory and it kept those copies in a central library even after its transformative use had been completed. But, unlike Texaco, which bought those copies, Anthropic never paid for the central library copies stolen off the internet. Texaco also shows why Anthropic is wrong to suppose that so long as you create an exciting end product, every “back-end step, invisible to the public,” is excused.
Notably, this is not a case where source copies were unavailable for separate purchase or loan. See, e.g., NXIVM Corp. v. Ross Inst., 364 F.3d 471, 475–76, 478–79 (2d Cir. 2004) (using selections of training manual — otherwise available only to cult’s trainees subject to NDAs — to expose cult in critical review); Time Inc. v. Bernard Geis Assocs., 293 F. Supp. 130, 135–36, 138, 146 (S.D.N.Y. 1968) (making charcoal drawings of photographs taken of originals otherwise not on sale or loan out to illustrate a history book). Nor were the copies made only incidentally and necessarily from pirated copies. See, e.g., Perfect 10, 508 F.3d at 1164 n.8 (copies of images that had been pirated by third-party websites were used to index those same websites while indexing the entire web). Here, piracy was the point: To build a central library that one could have paid for, just as Anthropic later did, but without paying for it.
Nor were the initial copies made immediately transformed into a significantly altered form. In Perfect 10, images were copied by the search engine in thumbnail form only and deployed immediately into the transformative use of identifying the full-sized images and the pages from which they came. 508 F.3d at 1160, 1165, 1167. And, in Kelly v. Arriba Software Corp., images were copied at full size and then into thumbnails for immediate use in building a search engine, after which the full-sized copies were immediately deleted. 336 F.3d 811, 815 (9th Cir. 2003). Not here. The full-text copies of books were downloaded and maintained “forever.”
Nor does the initial copying here even resemble the full-text copying in the Google Books cases. There, libraries of authorized copies already had been assembled, and all copies therefrom were made for direct employment in a one-to-one further fair use — whether the transformative use of pointing to the works themselves, the use of providing the works in formats for print-disabled patrons, or the use of insuring against going out of print, getting lost, and becoming otherwise unavailable. HathiTrust, 755 F.3d at 97, 101, 103; Google, 804 F.3d at 206, 216–18, 228 (further distinguishing search and snippet uses, which “test[ed] the boundaries of fair use”). Not so here concerning the pirated copies. No authorized copies existed from which Anthropic made its first copies. No full-text copy therefrom was put immediately into use training LLMs. Not every copy was even necessary nor used for training LLMs. No initial copy was ever deleted, even if never used or no longer used.7
Footnote 7: Training LLMs was not a use where perpetually maintaining a library copy was intrinsic to the proffered fair use (e.g., for a plagiarism-checker service). Nor is this an instance where retaining at least one copy was authorized by contract with the copyright owners (e.g., by agreement to express terms upon submission to a plagiarism-checker service, notwithstanding proposed terms scrawled on a paper prior to submission). A.V. ex rel. Vanderhye v. iParadigms, LLC, 562 F.3d 630, 635–36 & n.5, 645 n.8 (4th Cir. 2009), aff’g in relevant parts 544 F. Supp. 2d 473, 480 (E.D. Va. 2008) (Judge Claude Hilton). Anthropic mischaracterizes this case.
The university libraries and Google went to exceedingly great lengths to ensure that all copies were secured against unauthorized uses — both through technical measures and through legal agreements among all participants. Not so here. The library copies lacked internal controls limiting access and use.
Nor do the decisions on intermediate copying require anything less than the analysis applied here. Anthropic argues that our court of appeals in Sega Enterprises Ltd. v. Accolade, Inc. looked only at the “ultimate use” and “did not analyze a series of atomized acts of ‘infringement’ distinct from that overall purpose”. To the contrary, the appeals court examined the initial, intermediate, and ultimate copies used by the copyist. The court explained that the copyist initially purchased commercially available copies of game cartridges and then made further copies necessarily and “solely in order to discover the functional requirements for compatibility.” 977 F.2d 1510, 1522 (9th Cir. 1992). Thus, it reached only one result because on those facts there was only one “overall purpose” for the unauthorized copies. Indeed, the court reaffirmed prior caselaw holding that “intermediate copying of [a work] may infringe the exclusive rights granted to the copyright owner in [S]ection 106 of the Copyright Act regardless of whether the end product of the copying also infringes those rights.” Id. at 1518–19.
Similarly, in Sony Computer Entertainment, Inc. v. Connectix Corp., our appeals court applied the same law to similarly focused conduct. Another copyist allegedly had purchased an authorized copy and then made further copies solely and necessarily to reverse-engineer compatibility requirements. 203 F.3d 596, 601, 602–03 (9th Cir. 2000).
Both Sega and Sony avoided imposing an “artificial hurdle” to fair use by generously construing the intermediate copying necessary to the fair use. As one example, Sega stated that an engineer should be permitted to reboot her computer while undertaking to reverse-engineer software loaded onto it — even if doing so creates another digital copy of the software and is not strictly necessary to reverse-engineering. Id. at 605. But neither Sega nor Sony fathomed gifting an “artificial head start” to a fair user, either, by treating even the initial copy as an intermediate one.
And, yes, some courts have “not inquire[d]” into intermediate or initial copying at all (Reply 2 (citing Campbell as not inquiring into surplus copies in the studio)). But if a close reading of those cases reveals that in none of them was the legality of the initial or intermediate copying at issue, then it was not raised and not necessarily decided. It was expressly decided elsewhere: Our analysis must attend to different uses of different copies, and even to different uses of the same copies. Warhol, 598 U.S. at 533.
Finally, Anthropic argues that even if the initial copies served a different use than the intermediate and ultimate copies, it was not a use for which Anthropic necessarily would have needed to pay Authors for a copy. In theory, argues Anthropic, it could have done as Google did in Google Books — find an existing reference library willing to loan its copies for free as source copies. Or, in theory, it could have done as Anthropic did later — go buy used copies without having to pay Authors at all. See 17 U.S.C. § 109(a). But Anthropic did not do those things — instead it stole the works for its central library by downloading them from pirated libraries.
In sum, the first factor points against fair use for the central library copies made from pirated sources — and no damages from pirating copies could be undone by later paying for copies of the same works.
3. The Amount and Substantiality of the Portion Used.
B. The Copies Used to Build a Central Library.
But again, there was a separate use — a distinction that makes some difference as to whether the amount and substantiality of the copying was “reasonable in relation to the purpose of the copying” for the library copies. Campbell, 510 U.S. at 586.
(i) The Purchased Library Copies Converted from Print to Digital.
For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.
The third fair use factor favors fair use for the purchased library copies converted from print to digital.
(ii) The Pirated Library Copies.
For the pirated library copies, however, Anthropic lacked any entitlement to hold copies of the books at all. Its purpose, it says, was to train LLMs. But its objective conduct was to seek “all the books in the world” and then retain them even after deciding it would not make further copies from them for training — indicating there were other further uses. Against the purpose of acquiring all the books one could on the chance some might prove useful for training LLMs and maybe other stuff too, almost any unauthorized copying would have been too much. Anthropic copied millions of books in toto, Authors’ works among them.
The third factor points against fair use for the pirated library copies.
4. The Effect of the Use Upon the Market for or Value of the Copyrighted Work.
B. The Copies Used to Build A Central Library.
(i) The Purchased Library Copies Converted from Print to Digital.
For these copies, this order assumes Anthropic’s format change from print to digital displaced purchases of new digital copies that Anthropic would have made directly from Authors (had it not been able to purchase print copies in used condition). But for reasons stated under the first factor, such losses did not relate to something the Copyright Act reserves for Authors to exploit. It was a format change.
Authors’ next argument, it seems, is that the format change nonetheless exposed it to usurpation of the opportunity to sell rightful copies because Anthropic might transmit additional unauthorized digital copies more readily than it could have transmitted additional unauthorized print copies — and that the same would be true for all format converters. But after much discovery, there is no inkling in our record of intent to redistribute library copies once acquired nor of inability to secure that valuable library against outside actors. And, if the internal, central library copies did or do in fact lead to further reproduction or distribution, those further copies remain redressable separately by Authors. The format change did not itself usurp the Authors’ rightful entitlements.
This factor is thus neutral for the purchased library copies converted from print to digital.
(ii) The Pirated Library Copies.
The copies used to build a central library and that were obtained from pirated sources plainly displaced demand for Authors’ books — copy for copy. Not every person who merely intends to make a fair use of a work is thereby entitled to a full copy in the meantime, nor even to steal a copy so that achieving this fair use is especially simple or cost-effective. Here, the copies employed in training LLMs were one thing, but the copies acquired to assemble a convenient, general-purpose library of works for various uses for which the company might have of them, if any, was a different use altogether.
Anthropic has almost no rebuttal on these points. First, Anthropic argues that “Claude’s services do not reduce [or usurp] the value of Plaintiffs’ works through substitution in their traditional markets”. But stealing pirated copies of Authors’ works plainly did. Second, Anthropic argues that it may have been able to purchase some books on the open market (and some other texts), but not other texts it copied. But this case does not concern those other texts it could not have purchased. It could have purchased Authors’ books (and many others). In fact it later did. Finally, Anthropic argues that the effect on these texts from one book foregone was too small to be considered. But the test requires that we contemplate the likely result were the conduct to be condoned as a fair use — namely to steal a work you could otherwise buy (a book, millions of books) so long as you at least loosely intend to make further copies for a purportedly transformative use (writing a book review with excerpts, training LLMs, etc.), without any accountability. As Anthropic itself suggested, “That would destroy the [entire] publishing market if that were the case”.
The fourth factor points against fair use for the pirated library copies.
5. Overall Analysis.
The copies used to convert purchased print library copies into digital library copies were justified, too, though for a different fair use. The first factor strongly favors this result, and the third favors it, too. The fourth is neutral. Only the second slightly disfavors it. On balance, as the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.
The downloaded pirated copies used to build a central library were not justified by a fair use. Every factor points against fair use. Anthropic employees said copies of works (pirated ones, too) would be retained “forever” for “general purpose” even after Anthropic determined they would never be used for training LLMs. A separate justification was required for each use. None is even offered here except for Anthropic’s pocketbook and convenience.
And, as for any copies made from central library copies but not used for training, this order does not grant summary judgment for Anthropic. On this record in this posture, the central library copies were retained even when no longer serving as sources for training copies, “hundreds of engineers” could access them to make copies for other uses, and engineers did make other copies. Anthropic has dodged discovery on these points. We cannot determine the right answer concerning such copies because the record is too poorly developed as to them. Anthropic is not entitled to an order blessing all copying “that Anthropic has ever made after obtaining the data,” to use its words.
CONCLUSION
With respect to the training copies and the print-to-digital converted copies, this order has drawn all ambiguities and inferences in favor of the opposing side, namely Authors. With respect to the pirated copies, this order has also accepted the Authors’ version of the facts. Authors did not move for summary judgment but if they had, then we would have been obligated to accept all reasonable views given the evidence in defendant’s favor instead.
And, it grants that the print-to-digital format change was a fair use for a different reason. But it denies summary judgment for Anthropic that the pirated library copies must be treated as training copies.
We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages, actual or statutory (including for willfulness). That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for the theft but it may affect the extent of statutory damages. Nothing is foreclosed as to any other copies flowing from library copies for uses other than for training LLMs.
Notes and questions
(1) Notice how this decision applies Warhol’s “multiple uses” framework. Does the court’s distinction between “library building” and “training” uses seem appropriately granular, or could it lead to over-atomization of what should be considered unified activities?
(2) This decision creates a framework where the same copies can be treated differently based on their source and intended use. What might this mean for other AI companies’ training data acquisition strategies? Will this ruling help encourage the development of licensing markets for AI training? How will it impact the cost and complexity of developing competitive AI systems?
(3) Anthropic’s decision to acquire pirated copies first, then purchase legitimate copies later, seems to have significantly damaged its position. Would the analysis have been different if Anthropic had purchased legitimate copies first, then supplemented with pirated copies of unavailable works? What if Anthropic had made good-faith efforts to license works before resorting to havens of piracy? What if Anthropic had immediately deleted pirated copies after use for LLM training?
(4) Bartz v. Anthropic is a very strong vindication of print-to-digital format shifting. What are the limits of this decision?
(5) Note that although Judge Alsup’s condemnation of the use of pirate sites was broad, his actual holding in the Anthropic case is narrow. The ruling against Anthropic was based on the fact that it collected pirated books for purposes well beyond transformative LLM training and was building a general purpose library with no internal usage controls. On those facts, it is hard to argue with Alsup’s view that creating a permanent library from pirated books was a separate use unsupported by any transformative justification. Nonetheless Judge Alsup’s broader view on the use of shadow libraries is quite clear.
(6) The settlement in Bartz received final approval in July 2026, with the plaintiffs’ fee award reduced, so Judge Alsup’s fair use rulings were never tested on appeal. That leaves the shadow library question open. How should it be resolved? On the one hand, the fact that there is no “fruit of the poisonous tree” doctrine in copyright law suggests that treating what would otherwise be fair as unfair, simply because it was made from an infringing source, would deprive the public of the benefits fair use is meant to enable. See Mark A. Lemley, The Fruit of the Poisonous Tree in IP Law, 103 Iowa Law Review 245 (2017). However, arguably, when commercial users bypass the market for access without a compelling reason, they undermine the economic incentives that copyright is designed to create. See Matthew Sag, Fairness and Fair Use in Generative AI, 92 Fordham Law Review 1887, 1917–20 (2024). If this is correct, it explains Judge Alsup’s conclusion that although copyright owners do not have a right to charge for transformative uses, they should have a right to charge for access to their works.
A note on legislation
Congress has so far enacted nothing at all on AI training. Several transparency bills are pending: the TRAIN Act, reintroduced on 22 January 2026, would let a copyright owner obtain a clerk-issued subpoena compelling an AI developer to disclose whether particular works were used in training; a CLEAR Act introduced on 17 March 2026 and an AI Foundation Model Transparency Act introduced as H.R. 8094 would impose disclosure obligations directly. None has passed. The states have moved faster. California’s AB 2013 has been in force since 1 January 2026 and requires developers of generative AI systems made publicly available in California to post documentation of the datasets used to train them, including whether those datasets contain copyrighted material. It survived its first constitutional challenge, brought by xAI, in federal district court.
Selected overseas developments
Far too much is happening on the generative AI and copyright litigation front overseas to attempt an adequate summary here. This section simply calls out some of the most consequential developments. For a broader discussion of how copyright law applies to AI training in jurisdictions outside the United States, see Matthew Sag and Peter K. Yu, The Globalization of Copyright Exceptions for AI Training, 74 Emory Law Journal 1163 (2025).
Getty Images (US) Inc. v. Stability AI Ltd. [2025] EWHC 2863 (Ch) (4 Nov. 2025)
In the first UK judgment on generative AI training, Getty abandoned its primary copyright and database-right claims mid-trial, there being no evidence that training had occurred in the UK, and lost its secondary infringement claim that model weights are an “infringing copy.” It won only extremely limited trade mark findings on early versions of Stable Diffusion. Note that the court did not hold that training is lawful in the UK; the case turned on where the training took place, not on whether it was permissible.
Getty obtained permission to appeal the secondary infringement point on 16 December 2025. In the parallel US case (N.D. Cal. 3:25-cv-06891, Judge Trina Thompson), an order of 23 April 2026 let Getty’s trademark and unfair competition claims proceed and dismissed the DMCA false copyright management information claim for failure to plead scienter.
ANI Media v. Open AI (Delhi High Court, 24 July 2026)
In July 2026 the High Court in Delhi issued a decision on the legality of AI training and AI outputs under Indian copyright law in ANI Media v. Open AI. The result, on an interim injunction basis, is broadly congruent with Bartz and Kadrey, but the path to get there was different, because India does not have an open-ended fair use provision like the United States. Instead, Section 52(1)(a) of the Copyright Act, 1957 is a closed list of enumerated “fair dealing” purposes: private or personal use including research, criticism or review, and reporting of current events.
Justice Bansal, denying ANI’s interim injunction, held that OpenAI’s storage of scraped news articles for training fit within a relevant purpose and was fair. Specifically, he found that AI training counted as “private research.” In so doing he distinguished “private” from “personal” and therefore held that “private” did not exclude corporate commercial actors: “The expression ‘personal’ may be confined to individual persons, however, the term ‘private’ would include other private entities, including private companies.” ¶ 208. Central to the “private” holding was that the training corpus sits in a closed environment, never exposed to any person in natural-language or tokenized form: “The said data is not publicly available to any human entity either for access or for download.” ¶ 211.
Justice Bansal also gave “research” an updating construction reaching machine learning, reasoning that research conducted by machines is still conducted at the behest of and for the benefit of humans: “[T]he acts of further research cannot be confined to acts of human being alone and the same would extend to machine learning as well.” ¶ 218. The court cited an important Canadian fair dealing decision, CCH Canadian, for the proposition that research is not confined to non-commercial contexts.
Turning to whether the use was fair, Justice Bansal focused on three issues: whether the use was confined to training; whether it caused competitive harm to the rightsholder; and whether it served the public interest. On the first, OpenAI’s documents showed the models were not trained to reproduce or communicate the training material, and ANI identified no other use to which its articles had been put: “This Court has not been given any instance where Open AI has used the literary works of ANI for any purposes other than for training.” ¶ 243. On the second, the court reasoned that a general-purpose assistant and a news-syndication business perform fundamentally different functions, that in the news context ChatGPT supplies summaries or snippets with attribution and a link back to the source rather than the articles themselves, and that ANI had offered nothing beyond bare averments to show lost market share or reduced subscription revenue: “[I]t cannot be said that responses produced by ChatGPT are substitutes for news articles published by ANI.” ¶ 246. On the third, the court pointed to the benefits flowing from trained LLMs across access to information, education, scientific research, software development, translation, and accessibility tools, and treated those benefits as a legitimate input into the fairness inquiry rather than a policy aside: “While the rights of copyright owners remain important, the societal benefits arising from scientific and technological research constitute a relevant consideration in assessing the fairness of a dealing.” ¶ 253.
The output claim was resolved on familiar grounds. The court held that ANI has no copyright in the facts its journalists report, only in their expression, and found the ChatGPT responses submitted by the plaintiff were not in fact substantially similar to the underlying articles. The memorization allegations failed on the record, ANI’s illustrations all postdating the relevant training cutoffs. Every article ANI reproduced in its pleadings had been published in August or September 2024, whereas training had concluded in April 2022 for GPT-4 and April 2024 for GPT-4o. The court drew the obvious inference: material the model was never trained on cannot have been memorized from training, so the outputs must have been retrieved at query time. It described them as “in the nature of live links, perhaps reflecting RAG technique,” and observed that whether RAG outputs infringe was a distinct question ANI had not pleaded, though counsel raised it in argument. ¶ 84. That distinction also let the court set aside the Munich Regional Court’s contrary result in GEMA v. OpenAI, where verbatim song lyrics were reproduced in response to non-adversarial prompts with the search function disabled; here, by contrast, ANI’s prompts were repeated and detailed, at one point instructing the model to reproduce content “exactly,” and even then produced nothing substantially similar. ¶¶ 117–119.
One interesting aspect of the decision is that, when it came to assessing the merits of the interim injunction, the court’s balance-of-convenience discussion was unusually candid about industrial policy. The court observed that per-source licensing would make LLM development economically unviable and that an injunction would harm the growth of Indian models.
Further reading
On how copyright law should tackle the copying that takes place in the course of AI training, from different perspectives, see James Grimmelmann, Copyright for Literate Robots, 101 Iowa Law Review 657 (2016); Benjamin L. W. Sobel, Artificial Intelligence’s Fair Use Crisis, 41 Columbia Journal of Law & the Arts 45 (2017-2018); Matthew Sag, The New Legal Landscape for Text Mining and Machine Learning, 66 Journal of the Copyright Society of the U.S.A. 291 (2019); Mark Lemley and Bryan Casey, Fair Learning, 99 Texas Law Review 743 (2021); Peter Henderson et al., Foundation Models and Fair Use, 24 Journal of Machine Learning Research 1 (2023); Matthew Sag, Copyright Safety for Generative AI, 61 Houston Law Review 295 (2023); Pamela Samuelson, Fair Use Defenses In Disruptive Technology Cases, 71 UCLA Law Review 1484 (2024); David Opderbeck, Copyright in AI Training Data: A Human-Centered Approach, 76 Oklahoma Law Review 951 (2024); Robert Brauneis, Copyright and the Training of Human Authors and Generative Machines, 48 Columbia Journal of Law & the Arts 1 (2025).
Judge Leval is quite modest in not citing the Supreme Court’s extensive citation of the judge’s own law review article advocating transformativeness as the basis of fair use. See Pierre N. Leval, Toward a Fair Use Standard, 103 Harvard Law Review 1105 (1990).↩︎