Pillar 3 · Trustworthy self-improvement

AI Safety

Find dangerous behavior before deployment. We build adaptive attacks, agentic evaluations, and stress tests that reveal hidden failure modes across language models, generative systems, and deployed content safeguards.

Research question

How can we expose latent risks that ordinary evaluations and static red-teaming miss?

We build adaptive attacks, agentic evaluations, and stress tests that reveal hidden failure modes across language models, generative systems, and deployed content safeguards.

  • Agentic red-teaming
  • Adaptive attacks
  • Backdoor evaluation
  • Safety benchmarking

Representative projects

Find dangerous behavior before deployment.

Each project connects its central idea with papers, code, datasets, demonstrations, and public explanations.

Representative result from AdvBDGen
AI Safety2026

AdvBDGen

A robust framework for generating adaptive, stealthy backdoors that stress-test the resilience of LLM alignment defenses.

Representative result from AutoDAN
AI Safety2024

AutoDAN

Interpretable gradient-based adversarial prompts expose instruction-following vulnerabilities in aligned language models.

Representative result from WAVES
AI Safety2024

WAVES

A benchmark and red-team framework for measuring how image-watermark methods hold up under diverse attacks.

Selected publications

The papers behind the projects.

Browse the full publication database for the broader body of work in trustworthy ai.

Trustworthy AIconference2026

PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach

Udari Madhushani Sehwag, Shayan Shabihi, Alex McAvoy, Vikash Sehwag, Yuancheng Xu, Dalton Towers, Furong Huang

The Fourteenth International Conference on Learning Representations (ICLR), 2026

BibTeX
@inproceedings{sehwag2026propensitybench0e64,
  title = {PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach},
  author = {Udari Madhushani Sehwag and Shayan Shabihi and Alex McAvoy and Vikash Sehwag and Yuancheng Xu and Dalton Towers and Furong Huang},
  booktitle = {The Fourteenth International Conference on Learning Representations (ICLR), 2026},
  year = {2026},
  eprint = {2511.20703},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2511.20703},
}
Trustworthy AIconference2026

AdvBDGen: A Robust Framework for Generating Adaptive and Stealthy Backdoors in LLM Alignment Attacks

Pankayaraj Pathmanathan, Udari Madhushani Sehwag, Michael-Andrei Panaitescu-Liess, Cho-Yu Jason Chiang, Furong Huang

AAAI 2026 AI Alignment Track (AAAI), Oral, 2026

BibTeX
@inproceedings{pathmanathan2026advbdgen586f,
  title = {AdvBDGen: A Robust Framework for Generating Adaptive and Stealthy Backdoors in LLM Alignment Attacks},
  author = {Pankayaraj Pathmanathan and Udari Madhushani Sehwag and Michael-Andrei Panaitescu-Liess and Cho-Yu Jason Chiang and Furong Huang},
  booktitle = {AAAI 2026 AI Alignment Track (AAAI), Oral, 2026},
  year = {2026},
  eprint = {2410.11283},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2410.11283},
}
Trustworthy AIconference2024

AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, Tong Sun

First Conference on Language Modeling (COLM), 2024

BibTeX
@inproceedings{zhu2024autodan8585,
  title = {AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models},
  author = {Sicheng Zhu and Ruiyi Zhang and Bang An and Gang Wu and Joe Barrow and Zichao Wang and Furong Huang and Ani Nenkova and Tong Sun},
  booktitle = {First Conference on Language Modeling (COLM), 2024},
  year = {2024},
  eprint = {2310.15140},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2310.15140},
}
Trustworthy AIconference2024

WAVES: Benchmarking the Robustness of Image Watermarks

Bang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal, Yuancheng Xu, Chenghao Deng, Sicheng Zhu, Abdirisak Mohamed, Yuxin Wen, Tom Goldstein, Furong Huang

Proceedings of the 41st International Conference on Machine Learning (ICML), 2024

BibTeX
@inproceedings{an2024waves7a0c,
  title = {WAVES: Benchmarking the Robustness of Image Watermarks},
  author = {Bang An and Mucong Ding and Tahseen Rabbani and Aakriti Agrawal and Yuancheng Xu and Chenghao Deng and Sicheng Zhu and Abdirisak Mohamed and Yuxin Wen and Tom Goldstein and Furong Huang},
  booktitle = {Proceedings of the 41st International Conference on Machine Learning (ICML), 2024},
  year = {2024},
  eprint = {2401.08573},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2401.08573},
}