What Thomson Reuters v. Ross Does and Doesn’t Say NOW About Fair Use and Generative AI
Posted in AI
In February 2025, I assessed the operative summary judgment opinion in Thomson Reuters v. Ross and guessed at what it might mean for generative AI copyright litigation. I drew a distinction between the “small model” AI at issue in Ross and the large language models (LLMs) at issue in many of the generative AI copyright litigations. I conjectured that the decision would be instructive, but not determinative, in the generative AI cases.
A lot has happened since then. We have two subsequent decisions, Bartz v. Anthropic and Kadrey v. Meta, that found that the use of content to train LLMs is allowed under the fair use doctrine. And now the Third Circuit has affirmed that February 2025 opinion in Ross, adopting some aspects of Judge Stephanos Bibas’ opinion but rejecting others. Although the opinion attempted to distinguish the LLM cases, the analytical framework deployed in the opinion might impact future generative AI cases. In doing so, the Third Circuit has diverged from Bartz v. Anthropic, setting up a possible circuit split.
The Opinion Attempts to Distinguish Generative AI Cases, But …
The opinion noted that “ROSS’s AI was not a generative AI, meaning it would not create any new expression; it would only return text passages from preexisting judicial opinions.” On this basis, the opinion dismissed concerns raised by the Department of Justice (DOJ) in another copyright AI fair use case, In re OpenAI, Inc. Copyright Infringement Litigation, over restrictive fair use opinions and their impact on American competition in the artificial intelligence industry. The opinion further minimized the DOJ’s concerns by noting that LLM training may not result in “substitutive competition” with the works that they are trained on, while ROSS’s AI training was for the express purpose of competing with Thomson Reuters’ Westlaw product.
Thus, on the face of it, the opinion has limited applicability and ought not to have any impact on the LLM and other generative AI training cases. But the closer you look at the reasoning of the opinion, the more potentially impactful it becomes.
The Nature of the Transformation
The ROSS product was “an AI legal search engine that would respond to plain-language legal questions with relevant passages of text from judicial opinions.” To enable that search, the AI model “had to learn what made a judicial opinion responsive to a user’s legal question.” In other words, the “ROSS AI program [sought] to identify patterns regarding which judicial opinion passages respond well to legal questions.”
To do that, ROSS used a vendor to create legal memos that in some instances contained Westlaw headnotes, and it used the headnotes to benchmark and rank responses to legal queries. There was no evidence that the ROSS program ever outputted Westlaw headnotes or that the Westlaw headnotes were used after the AI model was trained.
Although the ROSS training was aimed at the noncopyrightable aspects of the Westlaw headnotes (i.e., detecting patterns of answer accuracy), the opinion nonetheless found the use was only minimally transformative: “ROSS’s platform uses Thomson Reuter’s headnotes to help users find judicial opinions related to their legal research inquiries, something Thomson Reuters already does with its headnotes.”
At a certain degree of generality (i.e., if you zoom out far enough), that makes some sense. It is true that the ROSS product as a whole (which used other answer-accuracy benchmarks, software development techniques and source code) generally competes with Westlaw as a whole (which contains other headnotes besides the ones used as well as other software features and types of content). But there is a difference between nontransformative substitutive use and commercial competition. As Judge William Alsup noted in Bartz, “The Act seeks to advance original works of authorship, not to protect authors against competition.”
Thus, the substitutive use must be aimed at the work itself (i.e., the headnotes used). This distinction was aptly articulated by Bibas in his 2023 summary judgment opinion: “[I]f Ross’s AI only studied the language patterns in the headnotes to learn how to produce judicial opinion quotes,” then ROSS’s use was transformative. If, however, ROSS’s use was “to get its AI to replicate and reproduce the creative drafting done by Westlaw’s attorney-editors,” then the use was not transformative. This distinction was lost in the subsequent 2025 summary judgment opinion and was not addressed at all in the Third Circuit opinion.
What Is ‘Necessary’?
Where the opinion stands to cause the most problems in future AI cases, however, is its application of a “strict” necessity standard to justify use of the underlying work. The opinion addressed the applicability of intermediate copying cases, in which software was downloaded to access noncopyrightable aspects of the software. The justification for copying in those cases was that there was no other way to access the nonprotectable aspects of the software but through the protectable code.
The opinion converts the factual background of those intermediate copying cases into an immutable maxim: “Here, ROSS does not need to copy Thomson Reuters’s headnotes to access the underlying unprotected information. … It chose not to do so because copying the headnotes offered an ‘easy’ way to create its training memos. Unlike necessity, ease is not a justification for copying.”
However, the standard for justifying use of the underlying work is not one of “strict necessity.” Under Warhol, the defendant need only show that the use was “reasonably necessary.” This distinction played an important role in Bartz, where the author plaintiffs argued that Anthropic did not strictly need to copy their specific books, or any books at all, to create its LLM. However, Anthropic demonstrated that there was a technological justification for the copying (i.e., the need to train on a “monumental” volume of texts), leading Alsup to conclude: “Because using so many works was reasonably necessary, using any one work for actually training LLMs was about as reasonable as the next.”
Where does one draw the line then between “ease” and “reasonable necessity”? What cost threshold did ROSS have to cross before using the Westlaw headnotes was easy versus necessary? What technological challenges had to be posed? We don’t know, because the opinion doesn’t say.
Impact on Future Cases
All eyes are currently on the generative AI copyright cases – 145 of them as of the date of this post. Those cases mostly involve LLMs or other multipurpose foundation models. But there is a great surge in enterprise use of AI to create fine-tuned or other “small models.” Those models might involve the use of content from competitors to obtain noncopyrightable information – market trends, competitive intelligence, product design ideas. If future courts interpret this opinion as prohibiting competitive uses and imposing a strict necessity standard, the impact it could have on the development of customized AI models is significant.
