Pillar 1 · World models

Multimodal learning

Connect perception, language, and generation. We build diagnostic benchmarks, structured datasets, and learning objectives that expose whether multimodal systems truly connect evidence across images, video, language, and generation.

Research question

How can a model reason across modalities instead of merely recognizing each one?

We build diagnostic benchmarks, structured datasets, and learning objectives that expose whether multimodal systems truly connect evidence across images, video, language, and generation.

  • Cross-modal reasoning
  • Sequential visual understanding
  • Diagnostic evaluation
  • Interleaved vision-language data

Representative projects

Connect perception, language, and generation.

Each project connects its central idea with papers, code, datasets, demonstrations, and public explanations.

Representative result from ROVER
Multimodal learning2025

ROVER

A benchmark for reciprocal reasoning between understanding and generation across image, video, audio, and 3D modalities.

Representative result from Mementos
Multimodal learning2024

Mementos

A comprehensive benchmark that tests whether multimodal language models can reason over coherent sequences of images.

Representative result from HallusionBench
Multimodal learning2024

HallusionBench

A diagnostic suite for disentangling visual illusion from language hallucination in large vision-language models.

Representative result from Zebra-CoT
Multimodal learning2025

Zebra-CoT

A dataset for teaching and evaluating interleaved vision-language chains of thought rather than text-only explanations.

Selected publications

The papers behind the projects.

Browse the full publication database for the broader body of work in world models.

World modelsconference2026

ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation

Yongyuan Liang, Wei Chow, Feng Li, Ziqiao Ma, Xiyao Wang, Jiageng Mao, Jiuhai Chen, Jiatao Gu, Yue Wang, Furong Huang

The Fourteenth International Conference on Learning Representations (ICLR), 2026

BibTeX
@inproceedings{liang2026roverbf75,
  title = {ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation},
  author = {Yongyuan Liang and Wei Chow and Feng Li and Ziqiao Ma and Xiyao Wang and Jiageng Mao and Jiuhai Chen and Jiatao Gu and Yue Wang and Furong Huang},
  booktitle = {The Fourteenth International Conference on Learning Representations (ICLR), 2026},
  year = {2026},
  eprint = {2511.01163},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2511.01163},
}
World modelsconference2024

Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Xiyao Wang, Yuhang Zhou, Xiaoyu Liu, Hongjin Lu, Yuancheng Xu, Feihong He, Jaehong Yoon, Taixi Lu, Fuxiao Liu, Gedas Bertasius, Mohit Bansal, Huaxiu Yao, Furong Huang

The 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024

BibTeX
@inproceedings{wang2024mementos5abe,
  title = {Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences},
  author = {Xiyao Wang and Yuhang Zhou and Xiaoyu Liu and Hongjin Lu and Yuancheng Xu and Feihong He and Jaehong Yoon and Taixi Lu and Fuxiao Liu and Gedas Bertasius and Mohit Bansal and Huaxiu Yao and Furong Huang},
  booktitle = {The 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024},
  year = {2024},
  eprint = {2401.10529},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2401.10529},
}
World modelsconference2024

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination & Visual Illusion in Large Vision-Language Models

Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, Tianyi Zhou

Conference on Computer Vision and Pattern Recognition (CVPR), 2024

BibTeX
@inproceedings{guan2024hallusionbench5094,
  title = {HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination \& Visual Illusion in Large Vision-Language Models},
  author = {Tianrui Guan and Fuxiao Liu and Xiyang Wu and Ruiqi Xian and Zongxia Li and Xiaoyu Liu and Xijun Wang and Lichang Chen and Furong Huang and Yaser Yacoob and Dinesh Manocha and Tianyi Zhou},
  booktitle = {Conference on Computer Vision and Pattern Recognition (CVPR), 2024},
  year = {2024},
  eprint = {2310.14566},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2310.14566},
}
World modelsconference2026

Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning

Ang Li, Charles L. Wang, Deqing Fu, Kaiyu Yue, Zikui Cai, Wang Bill Zhu, Ollie Liu, Peng Guo, Willie Neiswanger, Furong Huang, Tom Goldstein, Micah Goldblum

The Fourteenth International Conference on Learning Representations (ICLR), 2026

BibTeX
@inproceedings{li2026zebrab03e,
  title = {Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning},
  author = {Ang Li and Charles L. Wang and Deqing Fu and Kaiyu Yue and Zikui Cai and Wang Bill Zhu and Ollie Liu and Peng Guo and Willie Neiswanger and Furong Huang and Tom Goldstein and Micah Goldblum},
  booktitle = {The Fourteenth International Conference on Learning Representations (ICLR), 2026},
  year = {2026},
  eprint = {2507.16746},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2507.16746},
}