Skip to main content
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
  1. publications
  2. ai

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

Available Media Publication (PDF) GitHub
Conference Conference on Neural Information Processing Systems (NeurIPS) - 2026
Authors Zhun Wang , Nico Schiller , Hongwei Li ,
Citation BibTeX
BibTeX
@inproceedings{Wang2026ExploitGym,
  title = {ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?},
  author = {Zhun Wang and Nico Schiller and Hongwei Li and Srijiith Sesha Narayana and Milad Nasr and Nicholas Carlini and Xiangyu Qi and Eric Wallace and Elie Bursztein and Luca Invernizzi and Kurt Thomas and Yan Shoshitaishvili and Wenbo Guo and Jingxuan He and Thorsten Holz and Dawn Song},
  booktitle = {Conference on Neural Information Processing Systems},
  year = {2026},
  organization = {NeurIPS}
}

Finding a vulnerability and successfully exploiting it are different capabilities. ExploitGym evaluates the latter: starting from an input that triggers a bug, an AI agent must build a working exploit.

The benchmark contains 898 real-world instances spanning userspace applications, the V8 JavaScript engine and the Linux kernel. Reproducible environments and configurable defenses make it possible to study which protections limit agent success. The results provide a concrete way to measure the security implications of increasingly capable AI agents.

newsletter signup
newsletter signup