New study finds model outputs become unattributable as training data grows

MIT study: in large diffusion models, outputs can't be attributed to any single training image. Removing it doesn't change results, tested in 24 models up to 160,000 images.

Categorized in: AI News Science and Research
Published on: Aug 20, 2026
New study finds model outputs become unattributable as training data grows

Researchers at MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) have published a study showing that as generative AI models are trained on larger datasets, the connection between individual training examples and model outputs effectively disappears. The finding, published in Nature Communications, suggests that for large-scale diffusion models, it may be mathematically impossible to attribute any specific output to any specific piece of training data.

The team identified a phenomenon they call "attribution decay": the more data a model is trained on, the less any individual training example matters to any particular output. At sufficiently large scales, removing a single image - or every image by a given artist - does not change the generated result at all.

"If you take away a piece of data and the output of the model doesn't change, then that piece of data didn't affect the output," said Zheng Dai SM '21, PhD '24, former MIT CSAIL researcher and lead author on the work. "So it doesn't make much sense to attribute the output to that piece of data. And if you then do this one at a time for every other piece of data and find that the output doesn't change for any of them either, it doesn't make much sense to attribute the output to any one of them."

An exact method for testing attribution

Previous attempts to measure training data influence relied on approximations because testing directly requires retraining the model without each image - a prohibitive computation with millions of examples. The MIT team built a workaround: a "diffusion ensemble" architecture made of many smaller components, each trained on a different slice of the data. To test what the model would produce without a particular image, they switch off the parts that saw it. No retraining, no approximation.

"All previous methods were approximate," said MIT Professor David Gifford. "They could not absolutely show that deleting individual things did not change the output - but now, this paper introduces the first method that is absolute. You're actually deleting the inputs and deleting all influences of the inputs. This is the first exact method for doing large-scale deletion efficiently and showing that the results don't change."

The team trained 24 ensembles on datasets ranging from 256 images to over 160,000, drawn from seven public collections including CIFAR-10, CelebA, MetFaces, and ArtBench. The pattern held consistently: the larger the training set, the smaller the counterfactual radius - the maximum distance between an output and its most different alternate version with one training example removed. The decay followed an inverse power law and was statistically significant whether measured at the pixel level or by semantic meaning.

They stress-tested the result. At small scale, they trained 1,282 separate models the brute-force way and the decay still appeared. They pinned the removed fraction in place, fixed epochs, tried text-prompted and class-conditioned models, and used four similarity metrics. The finding survived every variation.

What this means for copyright and privacy

The implications cut directly into the legal debate over whether AI outputs are derivative works. If an output cannot be traced to any individual training example, the argument that it "copies" a specific artist or author becomes difficult to sustain.

"One way to think about this is that these models are creative. They are not simply the training data and echo data but with new outputs," Gifford said. "If those outputs have nothing to do with any individual training example, that raises the question about fair use, whether outputs are themselves copyrightable as novel works, and how authors get compensated when what comes out of a model is not attributable to anything on the internet."

Gifford also noted that the work provides a method for companies to prove their outputs are not derivative. "In order for the companies to show that their outputs aren't derivative of the internet in a copyright-infringing way, they need to revise their models to take advantage of the advances in this work, so they can prove they don't create derivatives of individual people or items."

James Grimmelmann, law professor at Cornell Law School and Cornell Tech, said the findings complicate existing legal frameworks. "If attribution worked, it would reliably show us that similarities between a model's output and a copyright-protected work are explicitly due to copying. But this paper provides reason to think attribution will fail for open and closed models. Instead, technologists and courts will need to resort to other methods for assessing protection."

The study examined diffusion models, which are widely used in audiovisual generation and scientific applications like protein structure modeling and therapeutic discovery. Whether the same decay holds for large language models - the technology at the center of the highest-profile copyright litigation - remains an open question.

Why this matters for science and research

For researchers building on generative models, the practical takeaway is direct: at scale, training data cannot be treated as individually accountable for outputs. This has consequences beyond copyright. Debugging a model by tracing an output to a training example becomes unreliable. Data removal requests - whether for privacy, licensing, or content moderation - may not change model behavior at all once a dataset crosses a certain size. Researchers working in AI for Science & Research should account for this when designing experiments that rely on model attribution, and when evaluating claims about what a model has learned.

The work also points to a practical method for guaranteeing unattributable outputs - a capability Gifford frames as an obligation rather than a loophole. For teams building Generative AI and LLM systems, the ensemble architecture offers a way to test counterfactual questions that were previously computationally out of reach.

The research was supported by Schmidt Futures.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)