Finding a vulnerability and successfully exploiting it are different capabilities. ExploitGym evaluates the latter: starting from an input that triggers a bug, an AI agent must build a working exploit.
The benchmark contains 898 real-world instances spanning userspace applications, the V8 JavaScript engine and the Linux kernel. Reproducible environments and configurable defenses make it possible to study which protections limit agent success. The results provide a concrete way to measure the security implications of increasingly capable AI agents.