Pillar 2 · Reasoning control

AI Agents

Turn models into purposeful systems. We develop agents that couple reasoning with action, critique, experimentation, and safety evaluation—especially where success requires long-horizon coordination.

Research question

How can an agent decide what to do, verify its progress, and recover when the first plan fails?

We develop agents that couple reasoning with action, critique, experimentation, and safety evaluation—especially where success requires long-horizon coordination.

  • Planning and verification
  • Tool use
  • Self-critique
  • Agentic evaluation

Representative projects

Turn models into purposeful systems.

Each project connects its central idea with papers, code, datasets, demonstrations, and public explanations.

Representative result from Agentic Critical Training
AI Agents2026

Agentic Critical Training

An agent-training framework that converts critique and revision into a learning signal for more reliable multi-step behavior.

Representative result from Imagine, Verify, Execute
AI Agents2025

Imagine, Verify, Execute

An embodied agent imagines candidate interactions, verifies which ones are informative, and executes a targeted exploration plan.

Representative result from SoundnessBench
AI Agents2026

SoundnessBench

An evaluation of whether AI scientist agents produce experiments and conclusions that are methodologically sound, not merely plausible.

Selected publications

The papers behind the projects.

Browse the full publication database for the broader body of work in reasoning control.

Reasoning controlpreprint2026

Agentic Critical Training

Weize Liu, Minghui Liu, Sy-Tuyen Ho, Souradip Chakraborty, Xiyao Wang, Furong Huang

BibTeX
@misc{liu2026agenticb95b,
  title = {Agentic Critical Training},
  author = {Weize Liu and Minghui Liu and Sy-Tuyen Ho and Souradip Chakraborty and Xiyao Wang and Furong Huang},
  year = {2026},
  eprint = {2603.08706},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2603.08706},
}
World modelsconference2025

Imagine, Verify, Execute: Agentic Exploration with Vision-Language Models

Seungjae^ Lee, Daniel Ekpo^, Haowen Liu, Furong Huang, Abhinav Shrivastava, Jia-Bin Huang

9th Annual Conference on Robot Learning (CoRL), 2025

BibTeX
@inproceedings{lee2025imagineef14,
  title = {Imagine, Verify, Execute: Agentic Exploration with Vision-Language Models},
  author = {Seungjae^ Lee and Daniel Ekpo^ and Haowen Liu and Furong Huang and Abhinav Shrivastava and Jia-Bin Huang},
  booktitle = {9th Annual Conference on Robot Learning (CoRL), 2025},
  year = {2025},
  eprint = {2505.07815},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2505.07815},
}
Trustworthy AIconference2026

PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach

Udari Madhushani Sehwag, Shayan Shabihi, Alex McAvoy, Vikash Sehwag, Yuancheng Xu, Dalton Towers, Furong Huang

The Fourteenth International Conference on Learning Representations (ICLR), 2026

BibTeX
@inproceedings{sehwag2026propensitybench0e64,
  title = {PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach},
  author = {Udari Madhushani Sehwag and Shayan Shabihi and Alex McAvoy and Vikash Sehwag and Yuancheng Xu and Dalton Towers and Furong Huang},
  booktitle = {The Fourteenth International Conference on Learning Representations (ICLR), 2026},
  year = {2026},
  eprint = {2511.20703},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2511.20703},
}