Task Sample · Posttrain

Learnability-Aware Long/Short CoT Adaptation

Adapt the supplied long and short mathematical reasoning examples so the same small language model learns more effectively under a fixed training recipe.

Best score per run5-task macro accuracy (pp) over AIME 2024, AMC 2023, OlympiadBench-EN, GSM8K test, MATH-500
051.3
Model · harnessBestSubmissionsRuntime
Claude Fable 5Claude Code47.5237.7 h
GPT-5.6 SolCodex · xhigh46.3191.79 h

Absolute five-benchmark macro accuracy for the candidate checkpoint; higher is better.

Teaser figure from Small-Model-Learnability-Gap (Small Models Struggle to Learn from Strong Reasoners)
Source project · Small-Model-Learnability-Gap · source ↗

The climb

All logged agent runs share one plot. Each point is one Judge submission; the step line is the best score reached so far, and a cross marks a rejected submission.

3438424650Baseline · 46.073 pp02467.7Elapsed time (h)5-task macro accuracyClaude Fable 5 · submission 1 · 46.1 · 0.29 h, Judge result recordedClaude Fable 5 · submission 2 · 0 · 0.789 h, Judge result recordedClaude Fable 5 · submission 3 · 0 · 1.71 h, Judge result recordedClaude Fable 5 · submission 4 · 44.2 · 2.19 h, Judge result recordedClaude Fable 5 · submission 5 · 0 · 2.42 h, Judge result recordedClaude Fable 5 · submission 6 · 43 · 2.7 h, Judge result recordedClaude Fable 5 · submission 7 · 36.4 · 3.1 h, Judge result recordedClaude Fable 5 · submission 8 · 44.2 · 3.38 h, Judge result recordedClaude Fable 5 · submission 9 · 44.2 · 3.67 h, Judge result recordedClaude Fable 5 · submission 10 · 45.5 · 4.01 h, Judge result recordedClaude Fable 5 · submission 11 · 43.9 · 4.28 h, Judge result recordedClaude Fable 5 · submission 12 · 45.4 · 4.62 h, Judge result recordedClaude Fable 5 · submission 13 · 41.1 · 4.93 h, Judge result recordedClaude Fable 5 · submission 14 · 44.7 · 5.19 h, Judge result recordedClaude Fable 5 · submission 15 · 43.3 · 5.45 h, Judge result recordedClaude Fable 5 · submission 16 · 42.3 · 5.72 h, Judge result recordedClaude Fable 5 · submission 17 · 42.2 · 5.99 h, Judge result recordedClaude Fable 5 · submission 18 · 44.6 · 6.24 h, Judge result recordedClaude Fable 5 · submission 19 · 44 · 6.49 h, Judge result recordedClaude Fable 5 · submission 20 · 47.5 · 6.75 h, Judge result recordedClaude Fable 5 · submission 21 · 42.6 · 7.15 h, Judge result recordedClaude Fable 5 · submission 22 · 44 · 7.41 h, Judge result recordedClaude Fable 5 · submission 23 · 45.7 · 7.67 h, Judge result recordedClaude Fable 5 · 47.5GPT-5.6 Sol · submission 1 · 0 · 0.207 h, Judge result recordedGPT-5.6 Sol · submission 2 · 0 · 0.228 h, Judge result recordedGPT-5.6 Sol · submission 3 · 0 · 0.241 h, Judge result recordedGPT-5.6 Sol · submission 4 · 0 · 0.248 h, Judge result recordedGPT-5.6 Sol · submission 5 · 0 · 0.256 h, Judge result recordedGPT-5.6 Sol · submission 6 · 0 · 0.281 h, Judge result recordedGPT-5.6 Sol · submission 7 · 0 · 0.291 h, Judge result recordedGPT-5.6 Sol · submission 8 · 0 · 0.296 h, Judge result recordedGPT-5.6 Sol · submission 9 · 0 · 0.321 h, Judge result recordedGPT-5.6 Sol · submission 10 · 0 · 0.347 h, Judge result recordedGPT-5.6 Sol · submission 11 · 0 · 0.352 h, Judge result recordedGPT-5.6 Sol · submission 12 · 0 · 0.382 h, Judge result recordedGPT-5.6 Sol · submission 13 · 0 · 0.398 h, Judge result recordedGPT-5.6 Sol · submission 14 · 0 · 0.405 h, Judge result recordedGPT-5.6 Sol · submission 15 · 0 · 0.413 h, Judge result recordedGPT-5.6 Sol · submission 16 · 0 · 0.447 h, Judge result recordedGPT-5.6 Sol · submission 17 · 45.1 · 0.801 h, Judge result recordedGPT-5.6 Sol · submission 18 · 44.9 · 1.31 h, Judge result recordedGPT-5.6 Sol · submission 19 · 46.3 · 1.76 h, Judge result recordedGPT-5.6 Sol · 46.3
The agent runs · 2 runs · submissionrejected submissionrunning bestbaseline · Judge measuredTime since run start · points mark recorded Judge results

The task

Adapt the supplied long and short mathematical reasoning responses into training data that a fixed small language model can learn from more effectively.

Environment

Reference baseline

The official long/short mixture checkpoint is evaluated under the same Judge protocol as the candidate. The page reports each candidate's absolute macro accuracy while the Judge also returns its difference from that baseline.

Research loop

  1. Inspect the reasoning pools, reference checkpoint, and earlier aggregate feedback.
  2. Form one hypothesis about what the small model can learn and materialize the adapted data.
  3. Train the fixed recipe, submit the complete checkpoint, and compare benchmark-level aggregates before revising.

What the agent may change

What stays fixed

Evaluation

The Judge evaluates the reference and candidate checkpoints on the same five mathematical-reasoning benchmarks. The score is the candidate's unweighted macro accuracy in percentage points; higher is better. Only aggregate benchmark results and training provenance are exposed.

Agent runs

Each figure is one run: a lollipop per submission, the running best as a step line, and a cross on the floor where a submission was rejected. Open the row below each figure for every Judge submission.

34.438.24245.749.547.5round 1235-task macro accuracyClaude Fable 5 · Claude Code
Claude Fable 5 · 23 submissions · best 47.5
Claude Fable 5 — every submission, in order (23)
#Score
146.1
20
30
444.2
50
643
736.4
844.2
944.2
1045.5
1143.9
1245.4
1341.1
1444.7
1543.3
1642.3
1742.2
1844.6
1944
2047.5
2142.6
2244
2345.7
44.645.145.64646.546.3round 1195-task macro accuracyGPT-5.6 Sol · Codex
GPT-5.6 Sol · 19 submissions · best 46.3
GPT-5.6 Sol — every submission, in order (19)
#Score
10
20
30
40
50
60
70
80
90
100
110
120
130
140
150
160
1745.1
1844.9
1946.3
Claude Fable 5 — the agent’s own closing summary

The best valid submission is round 24 / submission 20: macro 47.52, beating the official Mix-Long checkpoint's fixed-protocol score of 46.07. I verified the complete dataset, manifest, training records, tokenizer, and checkpoint artifacts.

The winning recipe was compress-then-select: trim verification tails from Long responses while preserving the final answer, then use the shortest compressed Long responses for a subset of problems and Short responses for the rest.

GPT-5.6 Sol — the agent’s own closing summary

Final audit passed: 5,383 records, 1,077 Long / 4,306 Short, zero rewrites, two completed epochs, canonical Qwen2 tokenizer and configuration, a 434-tensor safetensors topology, a reloadable checkpoint, and idle GPUs.

Links