angelX · measured telemetry
136 repository-repair tasks (JS, Python, Rust, C++) · 600 s wall cap · evaluator traces, zero self-reporting
| Harness | DeepSeek V4.1 Flash thinking off |
GLM-5.3-Flash thinking low |
Total both models |
|---|---|---|---|
| angelX | 133 / 136 | 135 / 136 | 268 / 272 98.5% |
| OpenCode 1.18.31 | 132 / 136 | 133 / 136 | 265 / 272 97.4% |
| oh-my-pi 18.2.4 | 53 / 59 budget cap | 133 / 136 | 186 / 195 95.4% |
Solved attempts vs cumulative agent time
Every attempt placed at the moment it completed
Tokens generated and model calls made per task (cap at 2.5k)
Wall time per attempt in run order (medians dotted)
Cumulative input tokens: cached vs fresh
Share of input tokens served from provider cache
* Graded by task-native test suites on Prime Intellect evaluators (Verifiers v0.3.1) · polyglot-v1: 136 tasks (48 JS, 34 Python, 30 Rust, 24 C++) · 600 s wall cap · 2026-09-21