MIT Study Finds Bigger AI Models Obscure Training-Data Origins
Copyright disputes over generative AI have often turned on whether a model’s output can be traced to particular training works and whether that use amounts to infringement or fair use. That framework becomes harder to apply as models learn from vast datasets and recombine features across many sources, potentially weakening the evidentiary link between an individual copyrighted work and a generated image.
MIT researchers found that generative diffusion models undergo “attribution degradation” as training datasets grow, making outputs increasingly difficult to connect to specific source material. Even after all works by a particular artist were removed, large models could still produce images with a highly similar style. Legal experts said attribution methods may therefore fail at scale, adding pressure on courts and the AI industry to develop new standards for assessing copying and plagiarism.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →