Reviewing AI Model Training Data Rights: The Chain-of-Title Problem

By Laith Sarhan

Data Protection & Cybersecurity Business Transactions

Every AI model has a chain of title problem, whether or not anyone building or buying it has looked at it directly. The question is simple to ask and hard to answer: where did the training data actually come from, and does the model provider — or you, if you fine-tuned on your own data — actually have the rights needed to use it that way?

This piece is about how to review that chain, not how to negotiate the contract clauses that follow from it (that's covered in the Enterprise AI MSA Playbook's IP and licensing chapter). This is the diligence question that comes before the negotiation.

The Four Data Sources, and What Each One Actually Requires

Training data behind any model generally falls into one of four buckets, and each carries a different set of questions.

1. Licensed corpora. Data the model provider (or you) obtained under an explicit license from a rights holder — a publisher, a data aggregator, a research consortium.

2. Scraped or crawled web content. Data collected via automated crawling, without a direct license from each individual source.

3. Customer or user data. Data your own customers provided to your product, now being considered for use in training or fine-tuning a model.

4. Synthetic data. Data generated by another model, rather than collected from any original human-authored source.

The Regulatory Overlay: What's Actually Required, Not Just Advisable

How This Is Evolving Through the Courts

The legal framework for AI training data isn't being built by legislatures — it's being built by courts, case by case, on both sides of the border.

The doctrinal gap that matters for diligence

US fair use is open-ended: any purpose can qualify if the four-factor test is satisfied. Canada's fair dealing is a closed list: only research, private study, education, parody/satire, criticism/review, and news reporting qualify (Copyright Act, s. 29). While Canadian courts interpret these purposes "large and liberally," uses of copyrighted material for purposes outside the list are excluded from the exception entirely — and "training a commercial AI model" doesn't map cleanly onto any enumerated purpose. This means the same training activity that might survive a fair use defense in the US could fail a fair dealing analysis in Canada on threshold grounds alone, before the fairness factors are even weighed.

The three US decisions that define the current landscape

Case Court / Date Outcome What it turns on
Thomson Reuters v. Ross Intelligence D. Del., Feb 11, 2025 Not fair use Ross used Westlaw headnotes to build a competing legal research tool. Direct market substitution; not transformative. First AI training fair use ruling on the merits. Not a generative AI case.
Bartz v. Anthropic N.D. Cal., June 23, 2025 Fair use for training; not fair use for pirated copies Training on copyrighted books was "exceedingly transformative." But downloading pirated copies to build a "library" was not fair use — trial ordered on damages. Resulted in a US$1.5 billion settlement, believed to be the largest copyright settlement in history.
Kadrey v. Meta N.D. Cal., June 25, 2025 Fair use, but explicitly narrow Training on copyrighted (and allegedly pirated) books was highly transformative. But the court stated the ruling "does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful" — the plaintiffs failed to develop market-harm evidence. Different evidence could have changed the outcome.

The emerging pattern across these three cases is not "training is always fair use" or "training is never fair use." It's market harm as the determining factor: when AI outputs compete with or substitute for the market the copyright owner controls (Thomson Reuters), courts find infringement; when they don't (Bartz, Kadrey), fair use survives — at least on the training itself. And there's a clear piracy bright line: using pirated source copies creates exposure that fair use may not cure, regardless of how transformative the training is (Bartz's pirated-copies ruling; Anthropic's $1.5B settlement).

Both Bartz and Kadrey are on appeal, and the bellwether New York Times v. OpenAI case (S.D.N.Y.) — where the court denied OpenAI's motion to dismiss, finding the NYT plausibly alleged that ChatGPT outputs compete with NYT content — is expected to reach trial in late 2026 or early 2027. The final shape of US fair use for AI training will likely depend on these appellate and trial outcomes.

The Canadian cases to watch

Two Canadian lawsuits are testing the same questions, both still in early stages with no decisions as of this writing:

What this means for your diligence review

The practical implication is that chain-of-title diligence on AI training data cannot be a static checklist when the law is evolving case-by-case. Three things follow:

  1. Document the market-relationship analysis. If your model's outputs could substitute for the works it was trained on (same industry, same audience, same use case), that's a materially higher risk profile than a model whose outputs serve a completely different purpose. This is the factor courts are actually weighing — your diligence should too.
  2. Treat pirated sources as an absolute red line, regardless of how transformative the training might be. The Bartz court was willing to call training "exceedingly transformative" and still ruled against the pirated copies. If your provenance review finds pirated content in the training set, that's a finding that needs to be disclosed and remediated, not argued away.
  3. Monitor the docket, not just the statute book. The law that will determine your exposure in 2026–2027 is being made in ongoing cases, not in legislation that hasn't been introduced. A diligence review completed today should be revisited when the next major ruling lands — particularly the NYT v. OpenAI trial and the Bartz/Kadrey appeals.

The Contractual Artifacts That Actually Manage This Risk

Diligence findings only matter if they translate into contract terms. Three specific artifacts do the work:

Why This Matters Even If You Didn't Build the Model

If you're licensing a third-party model — which is the position most companies are actually in — you don't get to skip this analysis by pointing at your vendor. Your own customers, regulators, and acquirers in a future transaction will look at your use of the model's outputs, and "our vendor's terms said it was fine" is a weaker position than having actually reviewed what those terms cover and don't cover, particularly in an M&A due diligence context where the acquirer's counsel will ask this question directly and expect a specific answer, not a shrug toward the vendor's fine print.

Last updated: August 2026. This article is educational and does not constitute legal advice. The legal status of text-and-data-mining exceptions, AI Act implementation timelines, and Canadian federal AI legislation are all subject to active change; verify current status directly before relying on any specific claim above in a live matter.