
Pillar 3 · Trustworthy self-improvement
Robust learning
Learn reliably under attack and shift. We study robust objectives and adaptive defenses across reinforcement learning, vision-language models, generative content, and adversarially corrupted training data.
Research question
How can a learning system adapt when its data, environment, or adversary changes?
We study robust objectives and adaptive defenses across reinforcement learning, vision-language models, generative content, and adversarially corrupted training data.
- Adaptive defense
- Data-poisoning resilience
- Robust reinforcement learning
- Watermark robustness
Representative projects
Learn reliably under attack and shift.
Each project connects its central idea with papers, code, datasets, demonstrations, and public explanations.




Selected publications
The papers behind the projects.
Browse the full publication database for the broader body of work in trustworthy ai.
WAVES: Benchmarking the Robustness of Image Watermarks
Proceedings of the 41st International Conference on Machine Learning (ICML), 2024
BibTeX
@inproceedings{an2024waves7a0c,
title = {WAVES: Benchmarking the Robustness of Image Watermarks},
author = {Bang An and Mucong Ding and Tahseen Rabbani and Aakriti Agrawal and Yuancheng Xu and Chenghao Deng and Sicheng Zhu and Abdirisak Mohamed and Yuxin Wen and Tom Goldstein and Furong Huang},
booktitle = {Proceedings of the 41st International Conference on Machine Learning (ICML), 2024},
year = {2024},
eprint = {2401.08573},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2401.08573},
}AdvBDGen: A Robust Framework for Generating Adaptive and Stealthy Backdoors in LLM Alignment Attacks
AAAI 2026 AI Alignment Track (AAAI), Oral, 2026
BibTeX
@inproceedings{pathmanathan2026advbdgen586f,
title = {AdvBDGen: A Robust Framework for Generating Adaptive and Stealthy Backdoors in LLM Alignment Attacks},
author = {Pankayaraj Pathmanathan and Udari Madhushani Sehwag and Michael-Andrei Panaitescu-Liess and Cho-Yu Jason Chiang and Furong Huang},
booktitle = {AAAI 2026 AI Alignment Track (AAAI), Oral, 2026},
year = {2026},
eprint = {2410.11283},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2410.11283},
}Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
The Thirty-eighth Annual Conference on Neural Information Processing Systems (NeurIPS), 2024
BibTeX
@inproceedings{xu2024shadowcastaafc,
title = {Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models},
author = {Yuancheng Xu and Jiarui Yao and Manli Shu and Yanchao Sun and Zichu Wu and Ning Yu and Tom Goldstein and Furong Huang},
booktitle = {The Thirty-eighth Annual Conference on Neural Information Processing Systems (NeurIPS), 2024},
year = {2024},
eprint = {2402.06659},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2402.06659},
}Beyond Worst-case Attacks: Robust RL with Adaptive Defense via Non-dominated Policies
Spotlight. The Twelfth International Conference on Learning Representations (ICLR), 2024
BibTeX
@inproceedings{liu2024beyond0ee1,
title = {Beyond Worst-case Attacks: Robust RL with Adaptive Defense via Non-dominated Policies},
author = {Xiangyu Liu and Chenghao Deng and Yanchao Sun and Yongyuan Liang and Furong Huang},
booktitle = {Spotlight. The Twelfth International Conference on Learning Representations (ICLR), 2024},
year = {2024},
eprint = {2402.12673},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2402.12673},
}