Seminar

Ensuring Data Ownership in Generative Visual Models

July 8, 2025 · Jun-Yan Zhu · Carnegie Mellon University
Ensuring Data Ownership in Generative Visual Models

Large-scale generative visual models have made content creation as little effort as writing a short text description. However, these models are typically trained on enormous amounts of Internet data, often containing copyrighted material, licensed images, and personal photos. How can we remove these images if the creators decide to opt out? How can we properly compensate them if they choose to opt in?

In this talk, I first describe an efficient method for removing copyrighted materials, artistic styles of living artists, and memorized images from pretrained text-to-image models. I then discuss our data-attribution algorithms for assessing the influence of each training image on a generated sample. Collectively, we aim to enable creators to retain control over the ownership of training images.

Speaker

Jun-Yan Zhu is the Michael B. Donohue Assistant Professor of Computer Science and Robotics at CMU’s School of Computer Science. Prior to CMU he was a Research Scientist at Adobe Research and a postdoc at MIT CSAIL. He obtained his Ph.D. from UC Berkeley and B.E. from Tsinghua University, and is a recipient of the Samsung AI Researcher of the Year, the Packard Fellowship, the NSF CAREER Award, and the ACM SIGGRAPH Outstanding Doctoral Dissertation Award.