Reverse engineering benchmark Hosted by Columbia University
SRE-Bench
A realistic, contamination-free benchmark for agentic reverse engineering.
A few weeks ago, we released SRE-Bench: The Next Challenge for Agentic Cybersecurity, a realistic, contamination-free benchmark designed to evaluate one of the most challenging capabilities in agentic cybersecurity: reverse engineering complex binary software.
Since then, we have been working with OpenAI to evaluate GPT-6 Astra, its strongest model. The results were striking: the model achieves a nearly 100% solve rate on SRE-Bench, while the other models we evaluated remain below 60%. This marks a remarkable advance in agentic reverse-engineering capabilities and illustrates just how rapidly the frontier of AI-driven cybersecurity is progressing.
At the same time, the strength of these results encouraged us to look beyond the current benchmark. Through a series of new pilot evaluations developed with OpenAI, we began exploring more challenging settings and gained an early glimpse of where our reverse-engineering expertise may help push the frontier even further. These preliminary findings motivate a long-term, community-driven effort to continually challenge, evaluate, and advance agentic cybersecurity.
SRE-Bench is now also hosted on Vals AI, making it easier for the broader community to evaluate and track progress in this rapidly evolving area.
Teams and sponsors that make SRE-Bench possible
Loading.
Citation
@article{spence2026srebench,
title = {The Next Challenge for Agentic Cybersecurity: A Realistic,
Contamination-Free Reverse Engineering Benchmark},
author = {Spence, Jeremy and Assaderaghi, Nicholas and Zhu, Jinhao and
Ravi, Nikil and Popa, Raluca Ada and Wei, Guannan and
Ding, Yangruibo and Zhang, Zhuo},
journal = {arXiv preprint arXiv:2608.11469},
year = {2026}
}