Benchmark Reference

Documentation

Task design, evaluation protocol, result definitions, and resources.

01

Overview

CLE evaluates experience-driven learning in professional workflows. Its 120 tasks span 12 occupational domains and 89 professional roles. Each task forms a five-level ladder, for a total of 600 levels.

Figure 1. CLE at a glance: 12 domains, 89 professional roles, and completion and improvement results across 17 agent configurations.
Figure 1. CLE at a Glance. View Full-Size Figure ↗

02

Task Structure

Levels are linked through earlier artefacts, tools, decisions, or professional knowledge. Passing a level unlocks the next one. The working environment and agent’s harness state persist as the task progresses.

Tools & skills
Reuse procedures and tool-use knowledge in later assignments.
Data Dependency
Use information and artefacts from earlier levels as inputs.
Strategy
Apply effective approaches and avoid recurring mistakes.
User Context
Retain preferences and constraints from prior interactions.
Domain Knowledge
Apply professional rules and conventions to new requirements.

The environments include 23,981 input files, 77 file extensions, 19 stateful mock applications, and 1,024 APIs.

03

A Task from the Manuscript

The example below summarizes the paper’s PhD research workflow. The Tasks page provides five complete workflows with instructions and input files.

TASK SPOTLIGHT

A Research Reproduction Sprint

Research / Education / Science
LEVEL 01 / LITERATURE CURATION

Build a Shared Reading Queue.

Collect the papers scattered across Slack, email, a WeChat screenshot, and the existing research plan. Verify their bibliographic details, add them to the group's Notion reading queue, and email a summary.

SlackEmailNotionWeb
WHAT CARRIES FORWARD
A Verified Research Foundation

The shared reading queue, verified references, and a record of the tools and decisions used to collect them.

Example Deliverablenotes_t1.md

Condensed from the paper's five-level task example. The explorer describes the assignments; it does not execute a benchmark run.

04

Evaluation Protocol

An evaluation agent checks submitted artefacts and application state against 16,125 human-authored rubric criteria. A level passes only when every checklist item is satisfied.

Trial 1 · K = 1

One Submission

One submission per level, without diagnostic feedback.

Trial 2 · K = 2

Feedback and Revision

A failed first submission receives diagnostic feedback and one opportunity to revise.

Feedback identifies unmet rubric categories. During evaluation, the agent does not receive the rubrics or reference answers. The public task showcase is for inspection; exhausting the submission budget ends the task, and unreached levels count as failed.

Model parameters stay fixed in the reported experiments. Agents can retain and update their harness state across the task stream.

05

Metrics

Completion
Levels cleared divided by 600, reported as a percentage.
L1–L5
Share of the 120 tasks whose first k levels are all cleared. The denominator includes tasks that never reached level k.
Full-Task Completion
The L5 pass rate: tasks completed end to end, divided by 120.
Δ Completion
Trial 2 completion minus Trial 1 completion, in percentage points (pp).

Trial differences include both feedback and an additional submission. They do not isolate the causal effect of a memory mechanism or learning strategy.

View the Complete Leaderboard →

06

Resources & availability

The Tasks page includes five tasks with their original instructions and inspectable input files. Reference answers and evaluation materials require approved access.

07

Citation

Use the provisional manuscript citation below. Publication details will be updated when a final reference is available.

BibTeX
@misc{cle2026continual,
  title  = {Continual Learning Exam},
  author = {{CLE Team}},
  year   = {2026},
  note   = {Manuscript}
}