Text and data mining: three legislative answers
Lesson 3 of 5 in AI and Intellectual Property: Training Data, Outputs, and the Litigation Wave.
Outside the US, legislatures did not wait for judges. Three jurisdictions wrote statutory text-and-data-mining exceptions, and the differences between them now shape where models get trained and what "lawful training data" means globally.
The EU answered in the 2019 DSM Copyright Directive. Article 3 gives research organisations and cultural-heritage institutions an unconditional right to mine works they lawfully access — rightsholders cannot opt out. Article 4 extends mining to everyone, for any purpose (commercial AI training included) — but only where the rightsholder has not reserved its rights, and for online content that reservation must be machine-readable. That opt-out clause quietly became the fulcrum of global AI copyright: robots.txt entries, metadata flags, and emerging standards for rights reservation are now legal instruments, and a German court gave the regime its first workout in LAION v. Kneschke (Hamburg, 2024), holding a nonprofit’s dataset assembly covered by the research exception.
Japan went furthest, earliest. Article 30-4 of its Copyright Act (2018) permits using works for purposes "not for enjoyment" of the expression — which squarely covers machine learning, commercially, with no opt-out mechanism. The caveat doing ever more work: the exception fails where use would "unreasonably prejudice the interests of the copyright owner", and 2024 guidance from Japan’s Agency for Cultural Affairs read that limit to catch fine-tuning aimed at reproducing a specific creator’s style and some RAG-style output uses. "The machine-learning paradise" has fences — but it remains the world’s most permissive major regime.
The UK is the cautionary tale of TDM politics. Its exception (CDPA s29A) covers non-commercial research only. A 2022 plan for a broad commercial exception died within months under creative-industry fire; a 2024–25 consultation proposing an EU-style opt-out model triggered the same battle, with high-profile artist campaigns and parliamentary ping-pong over transparency amendments. As of this writing the UK still has no commercial TDM exception — commercial training in the UK needs a licence — and reform remains contested. Check current status before advising.
Read Article 53 as the EU’s answer to a hard enforcement problem: training happens anywhere, but placing a model on the EU market happens in the EU. Any GPAI provider selling into Europe must honour EU opt-outs and publish a training-content summary regardless of where training ran — the AI Act’s recitals say a provider should not gain an advantage by training outside the EU to lower copyright standards. TDM opt-outs thus acquired extraterritorial pull, enforced not by copyright courts but by the AI Office, with AI Act penalties behind them. The Getty UK result — territoriality gutting the main claims — is exactly the gap this design closes.
Can you lawfully mine this content for training?
Interactive decision tree — outcomes:
- US: no statute — argue fair use
There is no US TDM exception; lawfulness rides on the §107 factors from the previous lesson. Acquisition legality, market substitution, and output regurgitation are the live edges.
- Stop — lawful access is the entry ticket
Both DSM exceptions require lawful access. Pirated or circumvented sources forfeit the exception before the analysis starts — the same lesson Bartz taught under US law.
- Covered — DSM Article 3
Research organisations mining lawfully accessed works for scientific research are covered, and no opt-out can stop them. This is the exception the Hamburg court applied to LAION’s dataset work.
- Covered — DSM Article 4, conditionally
Commercial mining is lawful where no machine-readable reservation exists. Keep the crawl-time evidence: under AI Act Article 53 you must operate a copyright policy that identifies and honours reservations, and your training-content summary is public.
- Blocked — licence or drop the source
A valid reservation removes Article 4. Your choices: negotiate a licence, or exclude the source and document the exclusion. Ignoring reservations now risks AI Act enforcement on top of infringement.
- Covered — Japan Article 30-4
Non-enjoyment uses including commercial ML are permitted with no opt-out — the most permissive major regime. But models placed on the EU market still owe EU-law compliance under AI Act Article 53.
- Outside the exception
Where use targets the expressive value — cloning a creator’s style, serving the work’s expression through outputs — the 2024 Agency for Cultural Affairs guidance reads Article 30-4’s prejudice proviso against you.
- Covered — CDPA s29A
Non-commercial research TDM with lawful access is covered. Note the boundary: commercialising the resulting model can take you outside it.
- No UK exception — licence required
Commercial TDM has no UK statutory cover as of this writing; reform proposals keep failing under creative-industry pressure. Licence the data, mine elsewhere, or check whether the law has finally changed.
Tool: Global Governance Atlas — Compare how jurisdictions split on TDM, transparency, and AI liability in the global regimes atlas.
Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.