Pillar 1 · World models

Embodied AI

We build agents that learn the structure of interaction—not only pixels or text—so knowledge can transfer across tasks, scenes, cameras, and bodies.

Research question

How can an agent learn a world model that is useful for action?

Our approach combines reusable physical representations, explicit state, multimodal reasoning, and active verification. The goal is not merely to generate plausible futures, but to support reliable decisions in the real world.

Representative projects

From video to interaction.

Each project connects its central idea with the paper, code, demonstrations, and public explanations.

Representative result from Imagine, Verify, Execute
Embodied AI2025

Imagine, Verify, Execute

An agentic exploration framework in which vision-language models imagine candidate interactions, verify their value, and execute promising actions.

Representative result from Make-An-Agent
Embodied AI2024

Make-An-Agent

A behavior-prompted diffusion framework that generates generalizable robot policies for seen and unseen manipulation tasks.

Selected publications

The papers behind the projects.

The publications database presents the broader body of work on world models and embodied intelligence.

World modelspreprint2026

μ0: A Scalable 3D Interaction-Trace World Model

Seungjae Lee, Yoonkyo Jung, Jusuk Lee, Jonghun Shin, Amir Hossein Shahidzadeh, Yao-Chih Lee, H. Jin Kim, Jia-Bin Huang, Furong Huang

BibTeX
@misc{lee2026scalablef557,
  title = {μ0: A Scalable 3D Interaction-Trace World Model},
  author = {Seungjae Lee and Yoonkyo Jung and Jusuk Lee and Jonghun Shin and Amir Hossein Shahidzadeh and Yao-Chih Lee and H. Jin Kim and Jia-Bin Huang and Furong Huang},
  year = {2026},
  eprint = {2606.13769},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.13769},
}
World modelsconference2026

TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos

Seungjae Lee, Yoonkyo Jung, Inkook Chun, Yao-Chih Lee, Zikui Cai, Hongjia Huang, Aayush Talreja, Tan Dat Dao, Yongyuan Liang, Jia-Bin Huang, Furong Huang

Conference on Computer Vision and Pattern Recognition (CVPR), 2026

BibTeX
@inproceedings{lee2026tracegen970c,
  title = {TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos},
  author = {Seungjae Lee and Yoonkyo Jung and Inkook Chun and Yao-Chih Lee and Zikui Cai and Hongjia Huang and Aayush Talreja and Tan Dat Dao and Yongyuan Liang and Jia-Bin Huang and Furong Huang},
  booktitle = {Conference on Computer Vision and Pattern Recognition (CVPR), 2026},
  year = {2026},
  eprint = {2511.21690},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2511.21690},
}
World modelspreprint2026

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Jusuk Lee, Seungjae Lee, Jonghun Shin, Hoseong Jung, Sungha Kim, Daesol Cho, H. Jin Kim, Jia-Bin Huang, Furong Huang

BibTeX
@misc{lee2026dynaflip336f,
  title = {DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation},
  author = {Jusuk Lee and Seungjae Lee and Jonghun Shin and Hoseong Jung and Sungha Kim and Daesol Cho and H. Jin Kim and Jia-Bin Huang and Furong Huang},
  year = {2026},
  eprint = {2605.30350},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.30350},
}
World modelsconference2026

MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning

Yuanchen Ju, Yongyuan Liang, Yen-Jen Wang, Nandiraju Gireesh, Yuanliang Ju, Seungjae Lee, Qiao Gu, Elvis Hsieh, Furong Huang, Koushil Sreenath

The Fourteenth International Conference on Learning Representations (ICLR), Oral, 2026

BibTeX
@inproceedings{ju2026momagraphf0b4,
  title = {MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning},
  author = {Yuanchen Ju and Yongyuan Liang and Yen-Jen Wang and Nandiraju Gireesh and Yuanliang Ju and Seungjae Lee and Qiao Gu and Elvis Hsieh and Furong Huang and Koushil Sreenath},
  booktitle = {The Fourteenth International Conference on Learning Representations (ICLR), Oral, 2026},
  year = {2026},
  eprint = {2512.16909},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2512.16909},
}
World modelsconference2025

Imagine, Verify, Execute: Agentic Exploration with Vision-Language Models

Seungjae^ Lee, Daniel Ekpo^, Haowen Liu, Furong Huang, Abhinav Shrivastava, Jia-Bin Huang

9th Annual Conference on Robot Learning (CoRL), 2025

BibTeX
@inproceedings{lee2025imagineef14,
  title = {Imagine, Verify, Execute: Agentic Exploration with Vision-Language Models},
  author = {Seungjae^ Lee and Daniel Ekpo^ and Haowen Liu and Furong Huang and Abhinav Shrivastava and Jia-Bin Huang},
  booktitle = {9th Annual Conference on Robot Learning (CoRL), 2025},
  year = {2025},
  eprint = {2505.07815},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2505.07815},
}
World modelsconference2024

Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted Diffusion

Yongyuan Liang, Tingqiang Xu, Kaizhe Hu, Guangqi Jiang, Furong Huang, Huazhe Xu

The Thirty-eighth Annual Conference on Neural Information Processing Systems (NeurIPS), 2024

BibTeX
@inproceedings{liang2024makee784,
  title = {Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted Diffusion},
  author = {Yongyuan Liang and Tingqiang Xu and Kaizhe Hu and Guangqi Jiang and Furong Huang and Huazhe Xu},
  booktitle = {The Thirty-eighth Annual Conference on Neural Information Processing Systems (NeurIPS), 2024},
  year = {2024},
  eprint = {2407.10973},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2407.10973},
}
World modelsworkshop2026

HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

Zhi Wang, Botao He, Kelin Yu, Seungjae Lee, Ruohan Gao, Furong Huang, Yiannis Aloimonos

Workshop on Data-Centric Robotics: What Data Do Robots Really Need, RSS 2026

BibTeX
@inproceedings{wang2026humanegoa2a2,
  title = {HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos},
  author = {Zhi Wang and Botao He and Kelin Yu and Seungjae Lee and Ruohan Gao and Furong Huang and Yiannis Aloimonos},
  booktitle = {Workshop on Data-Centric Robotics: What Data Do Robots Really Need, RSS 2026},
  year = {2026},
  url = {https://humanego-ai.github.io/},
}