Pillar 3 · Trustworthy self-improvement

Robust learning

Learn reliably under attack and shift. We study robust objectives and adaptive defenses across reinforcement learning, vision-language models, generative content, and adversarially corrupted training data.

Research question

How can a learning system adapt when its data, environment, or adversary changes?

We study robust objectives and adaptive defenses across reinforcement learning, vision-language models, generative content, and adversarially corrupted training data.

  • Adaptive defense
  • Data-poisoning resilience
  • Robust reinforcement learning
  • Watermark robustness

Representative projects

Learn reliably under attack and shift.

Each project connects its central idea with papers, code, datasets, demonstrations, and public explanations.

Representative result from WAVES
Robust learning2024

WAVES

A standardized evaluation exposes where image watermarks remain detectable—and where adaptive attacks break them.

Representative result from AdvBDGen
Robust learning2026

AdvBDGen

Adaptive backdoor generation provides a stronger adversary for testing and improving the robustness of alignment defenses.

Representative result from Shadowcast
Robust learning2024

Shadowcast

A stealthy data-poisoning attack reveals how small, targeted corruptions can alter the behavior of vision-language models.

Representative result from Adaptive Robust RL
Robust learning2024

Adaptive Robust RL

Non-dominated policies let a reinforcement-learning agent adapt its defense to attacks beyond a single fixed worst case.

Selected publications

The papers behind the projects.

Browse the full publication database for the broader body of work in trustworthy ai.

Trustworthy AIconference2024

WAVES: Benchmarking the Robustness of Image Watermarks

Bang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal, Yuancheng Xu, Chenghao Deng, Sicheng Zhu, Abdirisak Mohamed, Yuxin Wen, Tom Goldstein, Furong Huang

Proceedings of the 41st International Conference on Machine Learning (ICML), 2024

BibTeX
@inproceedings{an2024waves7a0c,
  title = {WAVES: Benchmarking the Robustness of Image Watermarks},
  author = {Bang An and Mucong Ding and Tahseen Rabbani and Aakriti Agrawal and Yuancheng Xu and Chenghao Deng and Sicheng Zhu and Abdirisak Mohamed and Yuxin Wen and Tom Goldstein and Furong Huang},
  booktitle = {Proceedings of the 41st International Conference on Machine Learning (ICML), 2024},
  year = {2024},
  eprint = {2401.08573},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2401.08573},
}
Trustworthy AIconference2026

AdvBDGen: A Robust Framework for Generating Adaptive and Stealthy Backdoors in LLM Alignment Attacks

Pankayaraj Pathmanathan, Udari Madhushani Sehwag, Michael-Andrei Panaitescu-Liess, Cho-Yu Jason Chiang, Furong Huang

AAAI 2026 AI Alignment Track (AAAI), Oral, 2026

BibTeX
@inproceedings{pathmanathan2026advbdgen586f,
  title = {AdvBDGen: A Robust Framework for Generating Adaptive and Stealthy Backdoors in LLM Alignment Attacks},
  author = {Pankayaraj Pathmanathan and Udari Madhushani Sehwag and Michael-Andrei Panaitescu-Liess and Cho-Yu Jason Chiang and Furong Huang},
  booktitle = {AAAI 2026 AI Alignment Track (AAAI), Oral, 2026},
  year = {2026},
  eprint = {2410.11283},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2410.11283},
}
Trustworthy AIconference2024

Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models

Yuancheng Xu, Jiarui Yao, Manli Shu, Yanchao Sun, Zichu Wu, Ning Yu, Tom Goldstein, Furong Huang

The Thirty-eighth Annual Conference on Neural Information Processing Systems (NeurIPS), 2024

BibTeX
@inproceedings{xu2024shadowcastaafc,
  title = {Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models},
  author = {Yuancheng Xu and Jiarui Yao and Manli Shu and Yanchao Sun and Zichu Wu and Ning Yu and Tom Goldstein and Furong Huang},
  booktitle = {The Thirty-eighth Annual Conference on Neural Information Processing Systems (NeurIPS), 2024},
  year = {2024},
  eprint = {2402.06659},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2402.06659},
}
Trustworthy AIconference2024

Beyond Worst-case Attacks: Robust RL with Adaptive Defense via Non-dominated Policies

Xiangyu Liu, Chenghao Deng, Yanchao Sun, Yongyuan Liang, Furong Huang

Spotlight. The Twelfth International Conference on Learning Representations (ICLR), 2024

BibTeX
@inproceedings{liu2024beyond0ee1,
  title = {Beyond Worst-case Attacks: Robust RL with Adaptive Defense via Non-dominated Policies},
  author = {Xiangyu Liu and Chenghao Deng and Yanchao Sun and Yongyuan Liang and Furong Huang},
  booktitle = {Spotlight. The Twelfth International Conference on Learning Representations (ICLR), 2024},
  year = {2024},
  eprint = {2402.12673},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2402.12673},
}