Overall Completion
The fraction of all 600 levels cleared. A level must pass every rubric criterion to count as complete.
120 Tasks · 600 Levels
Evaluation Protocol ↗17 Configurations · 4 Harnesses · 11 Backbone Variants
Showing all 17 configurations in Trial 2.
| # | Model / Harness | Completion | L1 | L2 | L3 | L4 | L5 | Δ CompletionTrial 2 − Trial 1 |
|---|---|---|---|---|---|---|---|---|
| 01 | Claude Opus 5Claude Code | 47.67% | 85.83% | 56.67% | 39.17% | 33.33% | 23.33% | +5.50 pp |
| 02 | Claude Sonnet 5Claude Code | 40.17% | 74.17% | 48.33% | 34.17% | 25.00% | 19.17% | +13.17 pp |
| 03 | Gemini 3.7 FlashHermes-Agent | 33.67% | 62.50% | 40.83% | 30.00% | 19.17% | 15.83% | +8.33 pp |
| 04 | GPT-5.6-SolCodex | 33.17% | 65.00% | 38.33% | 26.67% | 18.33% | 17.50% | +5.33 pp |
| 05 | GPT-5.6-TerraCodex | 33.17% | 68.33% | 40.83% | 25.83% | 18.33% | 12.50% | +5.33 pp |
| 06 | Doubao-Seed-EvolvingHermes-Agent | 30.00% | 62.50% | 38.33% | 22.50% | 14.17% | 12.50% | +9.67 pp |
| 07 | DeepSeek V4 ProHermes-Agent | 27.33% | 55.83% | 33.33% | 20.00% | 14.17% | 13.33% | +5.50 pp |
| 08 | Hy3Hermes-Agent | 27.33% | 60.83% | 35.00% | 19.17% | 11.67% | 10.00% | +7.00 pp |
| 09 | Qwen3.8-MaxHermes-Agent | 26.67% | 74.17% | 34.17% | 16.67% | 5.83% | 2.50% | +1.00 pp |
| 10 | GPT-5.6-LunaCodex | 25.33% | 56.67% | 30.83% | 19.17% | 11.67% | 8.33% | +5.00 pp |
| 11 | GLM 5.3Hermes-Agent | 24.00% | 62.50% | 30.00% | 15.00% | 7.50% | 5.00% | +1.50 pp |
| 12 | Gemini 3.7 FlashOpenClaw | 24.00% | 59.17% | 32.50% | 15.83% | 6.67% | 5.83% | +5.33 pp |
| 13 | Doubao-Seed-EvolvingOpenClaw | 21.17% | 48.33% | 26.67% | 15.83% | 8.33% | 6.67% | +6.67 pp |
| 14 | Qwen3.8-MaxOpenClaw | 19.83% | 65.00% | 20.00% | 10.83% | 1.67% | 1.67% | +1.67 pp |
| 15 | Hy3OpenClaw | 14.00% | 37.50% | 12.50% | 7.50% | 6.67% | 5.83% | +4.17 pp |
| 16 | GLM 5.3OpenClaw | 11.50% | 27.50% | 14.17% | 9.17% | 4.17% | 2.50% | +5.33 pp |
| 17 | DeepSeek V4 ProOpenClaw | 10.00% | 30.00% | 6.67% | 5.83% | 4.17% | 3.33% | −1.33 pp |
Completion = levels cleared / 600. Lk = share of 120 tasks whose first k levels are all cleared. L5 is full-task completion. Δ is measured in percentage points (pp).
These are results from complete runs under different submission budgets. The difference includes both diagnostic feedback and an additional attempt; it does not isolate a causal effect of memory or learning.
The fraction of all 600 levels cleared. A level must pass every rubric criterion to count as complete.
L1–L5 report cumulative pass rates across 120 tasks. L5 measures how often an agent completes an entire task.
Δ compares two complete runs. Trial 2 adds both feedback and a second attempt, so the change does not isolate learning alone.