One Submission
One submission per level, without diagnostic feedback.
Benchmark Reference
Task design, evaluation protocol, result definitions, and resources.
01
CLE evaluates experience-driven learning in professional workflows. Its 120 tasks span 12 occupational domains and 89 professional roles. Each task forms a five-level ladder, for a total of 600 levels.
02
Levels are linked through earlier artefacts, tools, decisions, or professional knowledge. Passing a level unlocks the next one. The working environment and agent’s harness state persist as the task progresses.
The environments include 23,981 input files, 77 file extensions, 19 stateful mock applications, and 1,024 APIs.
03
The example below summarizes the paper’s PhD research workflow. The Tasks page provides five complete workflows with instructions and input files.
Collect the papers scattered across Slack, email, a WeChat screenshot, and the existing research plan. Verify their bibliographic details, add them to the group's Notion reading queue, and email a summary.
The shared reading queue, verified references, and a record of the tools and decisions used to collect them.
notes_t1.mdCondensed from the paper's five-level task example. The explorer describes the assignments; it does not execute a benchmark run.
04
An evaluation agent checks submitted artefacts and application state against 16,125 human-authored rubric criteria. A level passes only when every checklist item is satisfied.
One submission per level, without diagnostic feedback.
A failed first submission receives diagnostic feedback and one opportunity to revise.
Feedback identifies unmet rubric categories. During evaluation, the agent does not receive the rubrics or reference answers. The public task showcase is for inspection; exhausting the submission budget ends the task, and unreached levels count as failed.
Model parameters stay fixed in the reported experiments. Agents can retain and update their harness state across the task stream.
05
Trial differences include both feedback and an additional submission. They do not isolate the causal effect of a memory mechanism or learning strategy.
View the Complete Leaderboard →06
The Tasks page includes five tasks with their original instructions and inspectable input files. Reference answers and evaluation materials require approved access.
07
Use the provisional manuscript citation below. Publication details will be updated when a final reference is available.
@misc{cle2026continual,
title = {Continual Learning Exam},
author = {{CLE Team}},
year = {2026},
note = {Manuscript}
}