Seminar

Recent Results on Multimodal Foundation Models

June 27, 2025 · Ming-Hsuan Yang · UC Merced · Google DeepMind
Recent Results on Multimodal Foundation Models

Recent advances in vision and language models have significantly improved visual understanding and generation tasks. In this talk, I present our latest research on designing effective tokenizers for transformers and our efforts to adapt frozen large language models for diverse vision tasks — including visual classification, video-text retrieval, visual captioning, visual question answering, visual grounding, video generation, stylization, outpainting, and video-to-audio conversion. Time permitting, I also discuss our recent findings on learning diffusion models and dynamic 3D vision.

Speaker

Ming-Hsuan Yang is a Professor at the University of California, Merced, and a Research Scientist at Google DeepMind. His awards include the Google Faculty Award, NSF CAREER Award, Nvidia Pioneer Research Award, the Longuet-Higgins (Test of Time) Prize at CVPR 2023, Best Paper at ICML 2024, and the WACV 2025 Test-of-Time Award. He is an Associate Editor-in-Chief of IEEE TPAMI and a Fellow of IEEE, ACM, and AAAI.