Leaderboard

We report two numbers for every model: full-solve rate, the share of binary instances on which the model completed every task, and score, the mean per-instance score. Rank by either one. Per-instance detail for each program and protection is on the Tasks page.

Rank by

Loading.

How a run is scored

We do not set a step limit. Each model runs until it decides it is done, and we grade what it leaves behind. Every binary instance carries several deterministically graded tasks, so grading is automatic and gives the same answer on every run.

Full-solve rate counts an instance only when every one of its tasks is completed. Score gives partial credit: it is the share of tasks completed on each instance, averaged over all instances. Switch between the two above the table.