Task Sample · Posttrain

ReasonIR Difficulty Curriculum

Choose how a fixed pool of easy and hard retrieval examples is weighted and ordered to improve ReasonIR-8B, without weakening its general retrieval ability.

Best score per runGuarded retrieval reward
00.241
Model · harnessBestSubmissionsRuntime
Claude Opus 5Claude Code · max0.223401024.01 h
GPT-5.6 SolCodex · xhigh0.21811824 h

Higher is better. The reward is the reasoning-retrieval score when general retrieval meets its 0.36126 reference; otherwise it is the negative shortfall from that reference. Negative values are measured regression penalties, not rejected submissions. Runtime is the reported agent wall time.

Teaser figure from ReasonIR
Source project · ReasonIR · source ↗

The climb

Both agent runs share one plot. Dots are Judge submissions and the step lines show each run's best score so far. The vertical scale is zoomed to make improvements readable. Hollow downward triangles mark the negative general-retrieval penalties below the displayed range; they are valid scores, not rejected submissions. Exact values remain in the tables below.

0.1960.2040.2120.220.228Baseline · 0.21141 reward081624.1Elapsed time (h)Guarded retrieval rewardClaude Opus 5 · submission 2 · 0.21776 · 5.36 h, Judge result recordedClaude Opus 5 · submission 3 · 0.21530 · 12.2 h, Judge result recordedClaude Opus 5 · submission 4 · 0.21161 · 12.8 h, Judge result recordedClaude Opus 5 · submission 5 · 0.21776 · 13.4 h, Judge result recordedClaude Opus 5 · submission 6 · 0.22340 · 17.8 h, Judge result recordedClaude Opus 5 · submission 8 · 0.22340 · 19.1 h, Judge result recordedClaude Opus 5 · submission 9 · 0.20748 · 21.6 h, Judge result recordedClaude Opus 5 · submission 10 · 0.22340 · 22.2 h, Judge result recordedClaude Opus 5 · 0.22340GPT-5.6 Sol · submission 1 · 0.20020 · 3.84 h, Judge result recordedGPT-5.6 Sol · submission 2 · 0.21536 · 6.4 h, Judge result recordedGPT-5.6 Sol · submission 5 · 0.21811 · 14.1 h, Judge result recordedGPT-5.6 Sol · submission 6 · 0.21260 · 16.8 h, Judge result recordedGPT-5.6 Sol · submission 7 · 0.21794 · 19.4 h, Judge result recordedGPT-5.6 Sol · 0.21811Claude Opus 5 · submission 1 · -0.00306 · 2.605 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 7 · -0.0049 · 18.421 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 3 · -0.01126 · 9.015 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 4 · -0.00368 · 11.532 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 8 · -0.01372 · 22.018 h, Judge result recorded; below the displayed range, marked at the axis.
The agent runs · 2 runs · submissionrunning bestbaseline · Judge measured▽ below 0.196Time since run start · points mark recorded Judge results

The task

Can a better training curriculum strengthen reasoning-intensive retrieval without sacrificing everyday retrieval? The agent starts from ReasonIR-8B and a fixed pool of hard-query (HQ) and varied-length-query (VL) examples. It changes which existing examples the model learns from, how often, and in what order. The source figure explains the original ReasonIR data recipe; this task uses its already-prepared HQ/VL pool and does not generate new data.

Environment

Reference baseline

The dashed 0.21141 baseline is the unchanged ReasonIR-8B checkpoint, submitted through the same adapter interface using an all-zero LoRA adapter. Its reasoning score is R0 = 0.21141 and its general retrieval score is G0 = 0.36126. An operator calibration and a separate formal Judge reproduced both values under this task's current protocol.

Every candidate is scored against these fixed references; Judge does not retrain or rerun the baseline beside it. Neither agent's first submission is the reference adapter. The original paper's full-BRIGHT result uses a different evaluation scope and is not this baseline.

Research loop

  1. Form a hypothesis about how example difficulty should affect learning.
  2. Define a policy over the fixed training pool and complete a matched-budget adapter-training run.
  3. Validate and submit the completed adapter with its manifest.
  4. Compare reasoning and general retrieval feedback, then retain or revise the policy.

A claim about why a curriculum works should be supported by matched policy trials, not just the best score. This reporting requirement is separate from Judge's scoring rules.

What the agent may change

What stays fixed

Evaluation

Judge validates the finished adapter and manifest, then evaluates the candidate without retraining it. Retrieval quality is measured with nDCG@10, which rewards placing relevant documents near the top of the first ten results.

Reasoning score R: the equal-weight mean over three BRIGHT subjects: Biology, Pony, and TheoremQA-theorems.

General retrieval score G: the equal-weight mean over four BEIR datasets: NFCorpus, SciFact, FiQA-2018, and ArguAna.

If G ≥ 0.36126, the reward is R. If general retrieval falls below that reference, the reward is G − 0.36126, a negative penalty. Higher is better in both cases. The means and reward use five decimal places.

The agent receives these aggregate scores, changes from the reference, and per-subject/dataset metrics. The evaluation uses the complete fixed query and candidate sets for the selected subjects; it is a three-subject reasoning proxy, not a full-BRIGHT score. Invalid or incomplete artifacts receive no score, distinct from a valid candidate's negative regression penalty.

Agent runs

Each figure shows one agent's submissions and running best. Hollow downward triangles mark negative general-retrieval penalties outside the zoomed range. Open the row beneath each figure for all submissions, including the exact negative scores.

0.1960.2040.2120.220.2280.22340Claude Opus 5 · Claude Code · submission 1 · -0.00306; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 7 · -0.0049; below the displayed range, marked at the axis.round 110Guarded retrieval rewardClaude Opus 5 · Claude Code
Claude Opus 5 · 10 submissions · best 0.22340 · ▽ below 0.196
Claude Opus 5 — every submission, in order (10)
#Score
1-0.00306
20.21776
30.21530
40.21161
50.21776
60.22340
7-0.00490
80.22340
90.20748
100.22340
0.1960.2040.2120.220.2280.21811GPT-5.6 Sol · Codex · submission 3 · -0.01126; below the displayed range, marked at the axis.GPT-5.6 Sol · Codex · submission 4 · -0.00368; below the displayed range, marked at the axis.GPT-5.6 Sol · Codex · submission 8 · -0.01372; below the displayed range, marked at the axis.round 18Guarded retrieval rewardGPT-5.6 Sol · Codex
GPT-5.6 Sol · 8 submissions · best 0.21811 · ▽ below 0.196
GPT-5.6 Sol — every submission, in order (8)
#Score
10.20020
20.21536
3-0.01126
4-0.00368
50.21811
60.21260
70.21794
8-0.01372

Links