Jonathan Hersh, PhD

Expert Witness Practice Area

AI Training Data & Copyright Expert Witness

In AI training data litigation, an economic expert quantifies whether a model's training and output caused market harm to the works it was trained on. Jonathan Hersh, PhD is an economist who analyzes substitution between AI outputs and original works, lost licensing markets, and the acquisition of training corpora — and has been deposed in an AI training data matter.

What these cases turn on

  • Did the model's outputs substitute for the original works, or complement them?
  • Was there a functioning licensing market for this training data, and what was it worth?
  • How was the training corpus acquired, and does the acquisition channel change the economic analysis?
  • Can observed harm to the rightsholder be separated from unrelated market trends?
  • What is the correct counterfactual — what would sales, licensing, or consumption have looked like absent the alleged conduct?

Analysis

AI training data disputes are, economically, questions about substitution. A rightsholder alleges that a model was built on their work and that the resulting outputs displaced demand for it. Answering that requires separating the effect of the alleged conduct from everything else moving in the market at the same time — the same identification problem that two decades of empirical work on digital copyright has been built around.

Why piracy economics applies to training data

Where a training corpus was assembled from unlicensed sources, the economic question is not new. It is the acquisition-and-displacement question that the piracy literature has studied with quasi-experimental methods since the early 2000s: when works are obtained outside the licensed channel, how much licensed demand is actually displaced, and how much would have existed anyway? My published research is in that literature — on website blocking, site shutdowns, and the measurable effect of enforcement on legal consumption. The estimation problem in an AI training data matter is structurally the same, and the same identification standards apply.

Where damages theories tend to fail

  • Assuming one-to-one displacement — that every AI output substitutes for a sale — without estimating the substitution rate.
  • Valuing a licensing market that did not exist and was not going to, or ignoring one that demonstrably did.
  • Attributing a revenue decline to the model when sector-wide trends explain it, with no control group or counterfactual.
  • Conflating the acquisition of training data with harm from model outputs; these are separate economic events and may carry different damages.
  • Relying on aggregate model capability claims rather than measured behavior on the works at issue.

My work in these matters is to build the counterfactual carefully, state the assumptions it rests on, and test whether the result survives when those assumptions are varied — then explain the result to a non-technical audience without softening what the data does and does not show.

Methods applied

  • Counterfactual demand modeling to separate alleged harm from background market trends
  • Difference-in-differences and synthetic control estimation on consumption and sales panels
  • Substitution and displacement analysis adapted from the empirical piracy literature
  • Licensing-market valuation where a comparable market exists or can be constructed
  • Model evaluation and output-similarity analysis using applied machine learning

Relevant engagements

  • Retained by the defense in a California matter concerning AI training data involving a Fortune 100 company. Provided economic analysis and was deposed. Further detail is limited by confidentiality.
  • Available for both plaintiff- and defense-side engagements. Conflicts check completed at intake.

Peer-reviewed research grounding this work

Published, citable research in this area — the verifiable basis for the analysis above.

Frequently asked questions

Have you testified in an AI training data case?

I was retained by the defense in a California matter concerning AI training data involving a Fortune 100 company, and was deposed in that engagement. I have not yet testified at trial. Further detail is limited by confidentiality obligations.

How do you quantify market harm from generative AI outputs?

By estimating substitution rather than assuming it. That means constructing a counterfactual for what demand would have been absent the alleged conduct, using difference-in-differences, synthetic control, or comparable quasi-experimental designs where the data supports them, and reporting how sensitive the estimate is to the assumptions behind it.

Does research on piracy actually apply to AI training data?

The identification problem is the same. Both ask how much licensed demand was displaced when works were obtained outside the licensed channel, and both require separating that effect from background market movement. The empirical piracy literature spent two decades developing methods for exactly this, and my peer-reviewed work is in it.

Can you work on either side of these matters?

Yes. The methods do not change with the retaining party, and I run a conflicts check before accepting any engagement.

Other practice areas

For expert witness matters or consulting, request a consult (conflicts check available).

Request a Consult