Reviewing AI Model Training Data Rights: The Chain-of-Title Problem
By Laith Sarhan
Data Protection & Cybersecurity Business Transactions
Every AI model has a chain of title problem, whether or not anyone building or buying it has looked at it directly. The question is simple to ask and hard to answer: where did the training data actually come from, and does the model provider — or you, if you fine-tuned on your own data — actually have the rights needed to use it that way?
This piece is about how to review that chain, not how to negotiate the contract clauses that follow from it (that's covered in the Enterprise AI MSA Playbook's IP and licensing chapter). This is the diligence question that comes before the negotiation.
The Four Data Sources, and What Each One Actually Requires
Training data behind any model generally falls into one of four buckets, and each carries a different set of questions.
1. Licensed corpora. Data the model provider (or you) obtained under an explicit license from a rights holder — a publisher, a data aggregator, a research consortium.
- What to check: the license's field-of-use restrictions (does it permit AI training specifically, or only some narrower use?), whether the license permits sublicensing model outputs to third parties, and whether the license terminates or restricts use if the underlying agreement lapses.
- The trap: a license negotiated before generative AI training was a foreseeable use often doesn't clearly cover it. A data license from 2015 that permits "internal research and analysis" is genuinely ambiguous about whether it covers training a commercial foundation model — and ambiguous licenses are exactly where disputes start.
2. Scraped or crawled web content. Data collected via automated crawling, without a direct license from each individual source.
- What to check: whether the crawl respected
robots.txtand any machine-readable opt-out signals, whether the source's terms of service prohibited automated collection (a contractual question, separate from copyright), and whether the jurisdiction where training occurred recognizes a text-and-data-mining exception that would cover the use even without a license. - The trap: this is the single most legally contested category right now, and the law is actively moving. In the EU, Article 4 of the Copyright in the Digital Single Market (DSM) Directive permits text-and-data-mining on lawfully accessible works unless the rights holder has reserved that right "in an appropriate manner" — and case law through 2025–2026 has been actively defining what "appropriate" means. A Danish court held in 2025 that a clear, accessible HTML policy statement (not necessarily a
robots.txtentry) was sufficient to constitute a valid reservation; a German court addressing a similar question required a machine-readable signal specifically. This is not settled law, and a pending CJEU referral (Like Company v. Google Ireland, C-250/25) is expected to bear directly on how "appropriate reservation" gets defined going forward. Treat any current summary of this area, including this one, as provisional. - Canada specifically: unlike the EU, Canada's Copyright Act does not currently contain an explicit text-and-data-mining exception for AI training. Whether existing fair dealing provisions (Copyright Act, s. 29) extend to AI training use is unresolved — ISED's 2023-24 consultation on copyright and generative AI surfaced sharp disagreement between technology-sector stakeholders (who argue existing exceptions already permit TDM) and rights-holder groups (who argue expressive content used in training is not merely "data" and that a TDM exception would fail the Berne Convention's three-step test for permissible exceptions). No legislative resolution has been enacted as of this writing.
3. Customer or user data. Data your own customers provided to your product, now being considered for use in training or fine-tuning a model.
- What to check: whether your terms of service or DPA with that customer actually permits this use — not just "improving the service," but specifically training or fine-tuning a model, ideally including whether outputs derived from it can be used with or sold to other customers. Under PIPEDA, if the data includes personal information, you need this to trace back to a valid basis under the "knowledge and consent" framework, and the OPC's guidelines on meaningful consent require that a reasonable person understand the nature, purpose, and consequences of the specific use — a generic "we may use your data to improve our services" clause is a real risk point if training use wasn't clearly within what a reasonable customer would have understood at the time of collection. Under GDPR, the equivalent question is which Article 6 legal basis applies (typically legitimate interest, requiring a documented balancing assessment, or consent).
- The trap: consent or contractual language obtained for one purpose (service delivery) doesn't automatically extend to a materially different purpose (training a model that benefits other customers, or that could be sold or licensed separately). This is the single most common gap found in due diligence on companies training on their own customer data.
4. Synthetic data. Data generated by another model, rather than collected from any original human-authored source.
- What to check: what training data produced the upstream model that generated your synthetic data, because synthetic data doesn't escape a chain-of-title problem — it inherits one, one level removed. If the upstream model was trained on improperly licensed data, synthetic outputs derived from it carry a version of the same exposure, even though no direct copy of the original work exists in your dataset.
- The trap: treating "it's synthetic, so there's no IP issue" as a settled legal conclusion. It isn't settled, and overclaiming that synthetic data is legally clean without having actually traced the upstream provenance is exactly the kind of unverifiable claim that creates liability later.
The Regulatory Overlay: What's Actually Required, Not Just Advisable
- EU AI Act, Article 53 requires providers of general-purpose AI models to maintain a copyright-compliance policy and to make publicly available a "sufficiently detailed summary" of the content used for training. This is a live obligation for GPAI providers placing models on the EU market, regardless of where the provider is headquartered.
- United States has no statutory text-and-data-mining exception and no federal AI training-data disclosure requirement. The sole defense for unauthorized use of copyrighted works in training is fair use under 17 U.S.C. § 107 — an open-ended, four-factor test (purpose and character of the use, nature of the copyrighted work, amount used, effect on the market) that applies to any purpose, not just a closed list. This is structurally different from Canada's fair dealing (limited to enumerated purposes in the Copyright Act) and from the EU's DSM Directive exception (which permits TDM subject to opt-out). The practical consequence: in the US, chain-of-title diligence for AI training is a fact-specific fair use risk assessment, not a compliance-box exercise — you're evaluating whether a court would likely find your specific use transformative and non-market-harming, not whether you satisfy a statutory exception. The US Copyright Office has issued guidance on AI output copyrightability (requiring human authorship for registration) but has not promulgated training-data-specific regulations.
- Canada has no equivalent federal AI-specific training-data disclosure requirement currently in force. Bill C-27, which contained the proposed Artificial Intelligence and Data Act (AIDA), died on the order paper when Parliament was prorogued in January 2025; as of this writing, a replacement federal AI framework has been publicly signaled but not introduced. Provincial measures (Quebec's Law 25, Ontario's Bill 194) address AI use in narrower contexts but don't fill this specific gap.
How This Is Evolving Through the Courts
The legal framework for AI training data isn't being built by legislatures — it's being built by courts, case by case, on both sides of the border.
The doctrinal gap that matters for diligence
US fair use is open-ended: any purpose can qualify if the four-factor test is satisfied. Canada's fair dealing is a closed list: only research, private study, education, parody/satire, criticism/review, and news reporting qualify (Copyright Act, s. 29). While Canadian courts interpret these purposes "large and liberally," uses of copyrighted material for purposes outside the list are excluded from the exception entirely — and "training a commercial AI model" doesn't map cleanly onto any enumerated purpose. This means the same training activity that might survive a fair use defense in the US could fail a fair dealing analysis in Canada on threshold grounds alone, before the fairness factors are even weighed.
The three US decisions that define the current landscape
| Case | Court / Date | Outcome | What it turns on |
|---|---|---|---|
| Thomson Reuters v. Ross Intelligence | D. Del., Feb 11, 2025 | Not fair use | Ross used Westlaw headnotes to build a competing legal research tool. Direct market substitution; not transformative. First AI training fair use ruling on the merits. Not a generative AI case. |
| Bartz v. Anthropic | N.D. Cal., June 23, 2025 | Fair use for training; not fair use for pirated copies | Training on copyrighted books was "exceedingly transformative." But downloading pirated copies to build a "library" was not fair use — trial ordered on damages. Resulted in a US$1.5 billion settlement, believed to be the largest copyright settlement in history. |
| Kadrey v. Meta | N.D. Cal., June 25, 2025 | Fair use, but explicitly narrow | Training on copyrighted (and allegedly pirated) books was highly transformative. But the court stated the ruling "does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful" — the plaintiffs failed to develop market-harm evidence. Different evidence could have changed the outcome. |
The emerging pattern across these three cases is not "training is always fair use" or "training is never fair use." It's market harm as the determining factor: when AI outputs compete with or substitute for the market the copyright owner controls (Thomson Reuters), courts find infringement; when they don't (Bartz, Kadrey), fair use survives — at least on the training itself. And there's a clear piracy bright line: using pirated source copies creates exposure that fair use may not cure, regardless of how transformative the training is (Bartz's pirated-copies ruling; Anthropic's $1.5B settlement).
Both Bartz and Kadrey are on appeal, and the bellwether New York Times v. OpenAI case (S.D.N.Y.) — where the court denied OpenAI's motion to dismiss, finding the NYT plausibly alleged that ChatGPT outputs compete with NYT content — is expected to reach trial in late 2026 or early 2027. The final shape of US fair use for AI training will likely depend on these appellate and trial outcomes.
The Canadian cases to watch
Two Canadian lawsuits are testing the same questions, both still in early stages with no decisions as of this writing:
- British Columbia proposed class action (MosaicML / Databricks, filed July 2025): alleges that MosaicML downloaded a dataset of pirated books to train its LLMs, and used software to remove copyright management information — a claim under the Copyright Act's anti-circumvention provisions as well as the core infringement claim.
- Ontario lawsuit (Canadian news media companies v. OpenAI): alleges that OpenAI improperly scraped copyrighted content from the plaintiffs' websites, reproduced it into training datasets, circumvented technological protection measures, and breached the terms of use governing the plaintiffs' websites.
What this means for your diligence review
The practical implication is that chain-of-title diligence on AI training data cannot be a static checklist when the law is evolving case-by-case. Three things follow:
- Document the market-relationship analysis. If your model's outputs could substitute for the works it was trained on (same industry, same audience, same use case), that's a materially higher risk profile than a model whose outputs serve a completely different purpose. This is the factor courts are actually weighing — your diligence should too.
- Treat pirated sources as an absolute red line, regardless of how transformative the training might be. The Bartz court was willing to call training "exceedingly transformative" and still ruled against the pirated copies. If your provenance review finds pirated content in the training set, that's a finding that needs to be disclosed and remediated, not argued away.
- Monitor the docket, not just the statute book. The law that will determine your exposure in 2026–2027 is being made in ongoing cases, not in legislation that hasn't been introduced. A diligence review completed today should be revisited when the next major ruling lands — particularly the NYT v. OpenAI trial and the Bartz/Kadrey appeals.
The Contractual Artifacts That Actually Manage This Risk
Diligence findings only matter if they translate into contract terms. Three specific artifacts do the work:
- A training-data provenance schedule attached to any AI vendor agreement or M&A data-asset transfer — a specific, factual description of the data sources, license basis, and any known gaps, rather than a general representation that "the model was trained in compliance with applicable law."
- A model card or equivalent technical disclosure, referenced by the contract, describing training data categories at a level of specificity that a counterparty's counsel can actually evaluate — vague enough to protect trade secrets, specific enough to be a real disclosure rather than a formality.
- Indemnity carve-outs scoped to the actual risk — a blanket IP indemnity from an AI vendor is only as good as the vendor's actual ability to pay out on it, and a diligence-informed negotiation should scope indemnity caps and carve-outs to what the provenance review actually revealed, rather than accepting boilerplate language that sounds protective but wasn't informed by a real review of the underlying data.
Why This Matters Even If You Didn't Build the Model
If you're licensing a third-party model — which is the position most companies are actually in — you don't get to skip this analysis by pointing at your vendor. Your own customers, regulators, and acquirers in a future transaction will look at your use of the model's outputs, and "our vendor's terms said it was fine" is a weaker position than having actually reviewed what those terms cover and don't cover, particularly in an M&A due diligence context where the acquirer's counsel will ask this question directly and expect a specific answer, not a shrug toward the vendor's fine print.
Last updated: August 2026. This article is educational and does not constitute legal advice. The legal status of text-and-data-mining exceptions, AI Act implementation timelines, and Canadian federal AI legislation are all subject to active change; verify current status directly before relying on any specific claim above in a live matter.