
Multimodal learning
Connect perception, language, and generation. We build diagnostic benchmarks, structured datasets, and learning objectives that expose whether multimodal systems truly connect evidence across images, video, language, and generation.
Research question
How can a model reason across modalities instead of merely recognizing each one?
We build diagnostic benchmarks, structured datasets, and learning objectives that expose whether multimodal systems truly connect evidence across images, video, language, and generation.
- Cross-modal reasoning
- Sequential visual understanding
- Diagnostic evaluation
- Interleaved vision-language data
Representative projects
Connect perception, language, and generation.
Each project connects its central idea with papers, code, datasets, demonstrations, and public explanations.




Selected publications
The papers behind the projects.
Browse the full publication database for the broader body of work in world models.
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
The Fourteenth International Conference on Learning Representations (ICLR), 2026
BibTeX
@inproceedings{liang2026roverbf75,
title = {ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation},
author = {Yongyuan Liang and Wei Chow and Feng Li and Ziqiao Ma and Xiyao Wang and Jiageng Mao and Jiuhai Chen and Jiatao Gu and Yue Wang and Furong Huang},
booktitle = {The Fourteenth International Conference on Learning Representations (ICLR), 2026},
year = {2026},
eprint = {2511.01163},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2511.01163},
}Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
The 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024
BibTeX
@inproceedings{wang2024mementos5abe,
title = {Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences},
author = {Xiyao Wang and Yuhang Zhou and Xiaoyu Liu and Hongjin Lu and Yuancheng Xu and Feihong He and Jaehong Yoon and Taixi Lu and Fuxiao Liu and Gedas Bertasius and Mohit Bansal and Huaxiu Yao and Furong Huang},
booktitle = {The 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024},
year = {2024},
eprint = {2401.10529},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2401.10529},
}HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination & Visual Illusion in Large Vision-Language Models
Conference on Computer Vision and Pattern Recognition (CVPR), 2024
BibTeX
@inproceedings{guan2024hallusionbench5094,
title = {HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination \& Visual Illusion in Large Vision-Language Models},
author = {Tianrui Guan and Fuxiao Liu and Xiyang Wu and Ruiqi Xian and Zongxia Li and Xiaoyu Liu and Xijun Wang and Lichang Chen and Furong Huang and Yaser Yacoob and Dinesh Manocha and Tianyi Zhou},
booktitle = {Conference on Computer Vision and Pattern Recognition (CVPR), 2024},
year = {2024},
eprint = {2310.14566},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2310.14566},
}Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning
The Fourteenth International Conference on Learning Representations (ICLR), 2026
BibTeX
@inproceedings{li2026zebrab03e,
title = {Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning},
author = {Ang Li and Charles L. Wang and Deqing Fu and Kaiyu Yue and Zikui Cai and Wang Bill Zhu and Ollie Liu and Peng Guo and Willie Neiswanger and Furong Huang and Tom Goldstein and Micah Goldblum},
booktitle = {The Fourteenth International Conference on Learning Representations (ICLR), 2026},
year = {2026},
eprint = {2507.16746},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2507.16746},
}