Reverse engineering benchmark Hosted by Columbia University
SRE-Bench
A realistic, contamination-free benchmark for agentic reverse engineering.
A few weeks ago, we released SRE-Bench: The Next Challenge for Agentic Cybersecurity, a realistic, contamination-free benchmark designed to evaluate one of the most challenging capabilities in agentic cybersecurity: reverse engineering complex binary software.
Since then, we have been working with OpenAI to evaluate GPT-6 Astra, its strongest model. The results were striking: the model achieves a nearly 100% solve rate on SRE-Bench, while the other models we evaluated remain below 60%. This marks a remarkable advance in agentic reverse-engineering capabilities and illustrates just how rapidly the frontier of AI-driven cybersecurity is progressing.
At the same time, the strength of these results encouraged us to look beyond the current benchmark. Through a series of new pilot evaluations developed with OpenAI, we began exploring more challenging settings and gained an early glimpse of where our reverse-engineering expertise may help push the frontier even further. These preliminary findings motivate a long-term, community-driven effort to continually challenge, evaluate, and advance agentic cybersecurity.
SRE-Bench is now also hosted on Vals AI, making it easier for the broader community to evaluate and track progress in this rapidly evolving area.
The teams behind SRE-Bench
Loading.
Citation
@article{spence2026srebench,
title = {The Next Challenge for Agentic Cybersecurity: A Realistic,
Contamination-Free Reverse Engineering Benchmark},
author = {Spence, Jeremy and Assaderaghi, Nicholas and Zhu, Jinhao and
Ravi, Nikil and Popa, Raluca Ada and Wei, Guannan and
Ding, Yangruibo and Zhang, Zhuo},
journal = {arXiv preprint arXiv:2608.11469},
year = {2026}
}