SREGym Leaderboard
Comparing SRE agents across diagnosis, mitigation, and end-to-end incident resolution on SREGym. Ranked by E2E success rate, requiring both correct root-cause diagnosis and successful mitigation on the same run.
Benchmark
SREGym-Lite-1004 · 17 faultsView cohort
1 | Codex | GPT-6 Astra (max) | 96.1 | 100.0 | 96.1 | 173.1 | 287.6 | 719K |
2 | Codex | GPT-6 Astra (medium) | 100.0 | 90.2 | 90.2 | 92.0 | 138.4 | 377K |
3 | Codex | GPT-6.1 Sol (max) | 98.0 | 90.2 | 90.2 | 217.4 | 371.6 | 807K |
4 | Codex | GPT-6.1 Sol (medium) | 98.0 | 90.2 | 90.2 | 103.8 | 156.6 | 462K |
5 | Codex | GPT-5.6 Sol (max) | 94.1 | 82.4 | 76.5 | 225.1 | 409.7 | 1.54M |
6 | CloudThinker* | Claude Opus 5 | 94.1 | 80.4 | 76.5 | 517.8 | 759.1 | 1.56M |
7 | Claude Code | Claude Opus 5 | 90.2 | 78.4 | 70.6 | 241.8 | 466.2 | 1.86M |
8 | Codex | GPT-5.6 Terra (max) | 82.4 | 74.5 | 62.7 | 233.1 | 444.9 | 1.88M |
9 | Codex | GPT-5.6 Luna (max) | 84.3 | 74.5 | 60.8 | 306.0 | 517.7 | 2.97M |
10 | Codex | GPT-5.6 Sol (medium) | 72.5 | 64.7 | 49.0 | 117.6 | 295.0 | 868K |
11 | Claude Code | Claude Sonnet 5 | 54.9 | 62.7 | 45.1 | 308.6 | 505.2 | 3.37M |
12 | Claude Code | Claude Opus 4.8 | 58.8 | 54.9 | 41.2 | 360.2 | 553.7 | 1.85M |
* Third-party submissions. Results verified by the SREGym team.
Diag. Diagnosis success rate · Mit. Mitigation success rate · E2E End-to-end (both diagnosis and mitigation correct) · TTD Time-to-diagnose (seconds) · TTM Time-to-mitigate (seconds) · Tokens Mean token usage per run