Leaderboard

120 Tasks · 600 Levels

Evaluation Protocol ↗

Main Evaluation

17 Configurations · 4 Harnesses · 11 Backbone Variants

Download Results CSV

Showing all 17 configurations in Trial 2.

Trial 2 · diagnostic feedback and up to two submissions per level. Higher percentages indicate better performance.
#Model / HarnessCompletionL1L2L3L4L5Δ CompletionTrial 2 − Trial 1
01Claude Opus 5Claude Code
47.67%
85.83%56.67%39.17%33.33%23.33%+5.50 pp
02Claude Sonnet 5Claude Code
40.17%
74.17%48.33%34.17%25.00%19.17%+13.17 pp
03Gemini 3.7 FlashHermes-Agent
33.67%
62.50%40.83%30.00%19.17%15.83%+8.33 pp
04GPT-5.6-SolCodex
33.17%
65.00%38.33%26.67%18.33%17.50%+5.33 pp
05GPT-5.6-TerraCodex
33.17%
68.33%40.83%25.83%18.33%12.50%+5.33 pp
06Doubao-Seed-EvolvingHermes-Agent
30.00%
62.50%38.33%22.50%14.17%12.50%+9.67 pp
07DeepSeek V4 ProHermes-Agent
27.33%
55.83%33.33%20.00%14.17%13.33%+5.50 pp
08Hy3Hermes-Agent
27.33%
60.83%35.00%19.17%11.67%10.00%+7.00 pp
09Qwen3.8-MaxHermes-Agent
26.67%
74.17%34.17%16.67%5.83%2.50%+1.00 pp
10GPT-5.6-LunaCodex
25.33%
56.67%30.83%19.17%11.67%8.33%+5.00 pp
11GLM 5.3Hermes-Agent
24.00%
62.50%30.00%15.00%7.50%5.00%+1.50 pp
12Gemini 3.7 FlashOpenClaw
24.00%
59.17%32.50%15.83%6.67%5.83%+5.33 pp
13Doubao-Seed-EvolvingOpenClaw
21.17%
48.33%26.67%15.83%8.33%6.67%+6.67 pp
14Qwen3.8-MaxOpenClaw
19.83%
65.00%20.00%10.83%1.67%1.67%+1.67 pp
15Hy3OpenClaw
14.00%
37.50%12.50%7.50%6.67%5.83%+4.17 pp
16GLM 5.3OpenClaw
11.50%
27.50%14.17%9.17%4.17%2.50%+5.33 pp
17DeepSeek V4 ProOpenClaw
10.00%
30.00%6.67%5.83%4.17%3.33%−1.33 pp

Completion = levels cleared / 600. Lk = share of 120 tasks whose first k levels are all cleared. L5 is full-task completion. Δ is measured in percentage points (pp).

These are results from complete runs under different submission budgets. The difference includes both diagnostic feedback and an additional attempt; it does not isolate a causal effect of memory or learning.

Reading the Results

Overall Completion

The fraction of all 600 levels cleared. A level must pass every rubric criterion to count as complete.

Depth of Completion

L1–L5 report cumulative pass rates across 120 tasks. L5 measures how often an agent completes an entire task.

Feedback and Revision

Δ compares two complete runs. Trial 2 adds both feedback and a second attempt, so the change does not isolate learning alone.