Task Sample · Vision

MolmoWeb Interaction Context Allocation

Choose which interaction history, page details, and screenshots a frozen MolmoWeb model sees to improve its next browser-action prediction.

Best score per runMacro action score (0–1)
00.654
Model · harnessBestSubmissionsRuntime
Claude Opus 5Claude Code · max0.60573924.01 h
GPT-5.6 SolCodex · xhigh0.60534024.01 h

Absolute macro action score over the same 948 held-out replay steps; higher is better. This measures next-action prediction, not live-browser task completion. Counts include every persisted Judge submission; runtime is the reported agent wall time.

Teaser figure from MolmoWeb
Source project · MolmoWeb · source ↗

The climb

All logged agent runs share one plot. Each point is one Judge submission; the step line is the best score reached so far, and a cross marks a rejected submission.

0.5920.5960.60.6040.608Baseline · 0.5966 action score081624.3Elapsed time (h)Macro action scoreClaude Opus 5 · submission 1 · 0.5966 · 0.784 h, Judge result recordedClaude Opus 5 · submission 2 · 0.5983 · 1.41 h, Judge result recordedClaude Opus 5 · submission 5 · 0.5932 · 3.17 h, Judge result recordedClaude Opus 5 · submission 7 · 0.5965 · 4.4 h, Judge result recordedClaude Opus 5 · submission 8 · 0.6057 · 5 h, Judge result recordedClaude Opus 5 · submission 9 · 0.5997 · 5.6 h, Judge result recordedClaude Opus 5 · submission 10 · 0.5934 · 6.21 h, Judge result recordedClaude Opus 5 · submission 11 · 0.6036 · 6.81 h, Judge result recordedClaude Opus 5 · submission 12 · 0.6040 · 7.41 h, Judge result recordedClaude Opus 5 · submission 13 · 0.6033 · 8.03 h, Judge result recordedClaude Opus 5 · submission 14 · 0.5973 · 8.63 h, Judge result recordedClaude Opus 5 · submission 15 · 0.6057 · 9.23 h, Judge result recordedClaude Opus 5 · submission 16 · 0.6013 · 9.87 h, Judge result recordedClaude Opus 5 · submission 19 · 0.6057 · 11.7 h, Judge result recordedClaude Opus 5 · submission 20 · 0.6057 · 12.3 h, Judge result recordedClaude Opus 5 · submission 23 · 0.6057 · 14.1 h, Judge result recordedClaude Opus 5 · submission 26 · 0.6057 · 16 h, Judge result recordedClaude Opus 5 · submission 27 · 0.5959 · 16.6 h, Judge result recordedClaude Opus 5 · submission 28 · 0.6057 · 17.3 h, Judge result recordedClaude Opus 5 · submission 29 · 0.5993 · 18 h, Judge result recordedClaude Opus 5 · submission 30 · 0.6057 · 18.5 h, Judge result recordedClaude Opus 5 · submission 32 · 0.6057 · 19.8 h, Judge result recordedClaude Opus 5 · submission 34 · 0.6057 · 21 h, Judge result recordedClaude Opus 5 · submission 35 · 0.6044 · 21.6 h, Judge result recordedClaude Opus 5 · submission 37 · 0.6057 · 22.8 h, Judge result recordedClaude Opus 5 · submission 39 · 0.6057 · 24.1 h, Judge result recordedClaude Opus 5 · 0.6057GPT-5.6 Sol · submission 1 · 0.5966 · 0.58 h, Judge result recordedGPT-5.6 Sol · submission 2 · 0.5985 · 1.78 h, Judge result recordedGPT-5.6 Sol · submission 4 · 0.5940 · 3.01 h, Judge result recordedGPT-5.6 Sol · submission 5 · 0.5985 · 3.63 h, Judge result recordedGPT-5.6 Sol · submission 6 · 0.5975 · 4.22 h, Judge result recordedGPT-5.6 Sol · submission 7 · 0.5991 · 4.81 h, Judge result recordedGPT-5.6 Sol · submission 8 · 0.6011 · 5.4 h, Judge result recordedGPT-5.6 Sol · submission 9 · 0.5959 · 6.01 h, Judge result recordedGPT-5.6 Sol · submission 10 · 0.6001 · 6.68 h, Judge result recordedGPT-5.6 Sol · submission 11 · 0.6011 · 7.27 h, Judge result recordedGPT-5.6 Sol · submission 13 · 0.6009 · 8.45 h, Judge result recordedGPT-5.6 Sol · submission 14 · 0.6011 · 9.04 h, Judge result recordedGPT-5.6 Sol · submission 15 · 0.6020 · 9.65 h, Judge result recordedGPT-5.6 Sol · submission 16 · 0.6022 · 10.2 h, Judge result recordedGPT-5.6 Sol · submission 20 · 0.6020 · 12.3 h, Judge result recordedGPT-5.6 Sol · submission 21 · 0.6031 · 12.9 h, Judge result recordedGPT-5.6 Sol · submission 25 · 0.6044 · 15.3 h, Judge result recordedGPT-5.6 Sol · submission 26 · 0.6028 · 15.9 h, Judge result recordedGPT-5.6 Sol · submission 27 · 0.6044 · 16.5 h, Judge result recordedGPT-5.6 Sol · submission 29 · 0.6040 · 17.7 h, Judge result recordedGPT-5.6 Sol · submission 30 · 0.6049 · 18.3 h, Judge result recordedGPT-5.6 Sol · submission 31 · 0.6031 · 18.9 h, Judge result recordedGPT-5.6 Sol · submission 32 · 0.6049 · 19.5 h, Judge result recordedGPT-5.6 Sol · submission 33 · 0.5971 · 20.1 h, Judge result recordedGPT-5.6 Sol · submission 34 · 0.6043 · 20.7 h, Judge result recordedGPT-5.6 Sol · submission 35 · 0.5993 · 21.3 h, Judge result recordedGPT-5.6 Sol · submission 36 · 0.5995 · 21.9 h, Judge result recordedGPT-5.6 Sol · submission 37 · 0.6053 · 22.5 h, Judge result recordedGPT-5.6 Sol · submission 38 · 0.5988 · 23.1 h, Judge result recordedGPT-5.6 Sol · submission 39 · 0.5985 · 23.7 h, Judge result recordedGPT-5.6 Sol · submission 40 · 0.6039 · 24.3 h, Judge result recordedGPT-5.6 Sol · 0.6053Claude Opus 5 · submission 3 · 0.5624024003325582 · 2.176 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 4 · 0.31013111490596057 · 2.544 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 6 · 0.588914339543402 · 3.782 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 17 · 0.571584097878923 · 10.490 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 18 · 0.5721456164191883 · 11.101 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 21 · 0.5808311487459865 · 12.892 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 22 · 0.5616897313679421 · 13.529 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 24 · 0.47783659480152735 · 14.775 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 25 · 0.2821666630983059 · 15.420 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 31 · 0.49152040367063765 · 19.211 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 33 · 0.5668735513918998 · 20.447 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 36 · 0.5866893299751684 · 22.246 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 38 · 0.5769691386941643 · 23.455 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 3 · 0.5601990639390262 · 2.408 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 12 · 0.5905210272412961 · 7.866 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 17 · 0.5735986250297 · 10.799 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 18 · 0.5533460410634136 · 11.393 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 19 · 0.3039013154057625 · 11.724 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 22 · 0.48974277196778726 · 13.490 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 23 · 0.5574342740981484 · 14.121 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 24 · 0.5887556372220799 · 14.712 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 28 · 0.5740402176998849 · 17.108 h, Judge result recorded; below the displayed range, marked at the axis.
The agent runs · 2 runs · submissionrunning bestbaseline · Judge measured▽ below 0.592Time since run start · points mark recorded Judge results

The task

Find a better mix of textual history, page state, and visual context for a frozen MolmoWeb-4B model. The model predicts the next action from recorded browser steps; it is not trained or evaluated by completing live browsing tasks.

Environment

Reference baseline

The shipped strategy uses recent textual thought/action history and the current screenshot, with no past screenshots. Under the fixed greedy replay protocol it scored 0.5966, reproduced by the first submission of both runs.

Every candidate is evaluated on the same held-out steps and receives an absolute score; a separate baseline is not rerun beside each candidate. The source project's live-browser success rates are not comparable to this offline replay score.

Research loop

  1. Form a hypothesis about which past interactions or page details help predict the next action.
  2. Edit the strategy and validate its schema and context budget.
  3. Optionally test on the public proxy trajectories, then submit the closed strategy file.
  4. Compare aggregate feedback by action type and refine the context choices.

What the agent may change

What stays fixed

Evaluation

The Judge evaluates all 948 held-out replay steps. It first averages graded prediction scores within each reference action type, then averages those action-type means equally. The resulting score lies between zero and one; higher is better.

Feedback includes overall and per-action aggregates, parsing validity, and context usage. Hidden trajectories, targets, predictions, and individual outcomes remain concealed. A strategy that violates the file, schema, or context constraints receives zero; an infrastructure failure receives no score.

Agent runs

Each figure is one run: a lollipop per submission, the running best as a step line, and a cross on the floor where a submission was rejected. Open the row below each figure for every Judge submission.

0.5920.5960.60.6040.6080.6057Claude Opus 5 · Claude Code · submission 3 · 0.5624024003325582; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 4 · 0.31013111490596057; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 6 · 0.588914339543402; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 17 · 0.571584097878923; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 18 · 0.5721456164191883; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 21 · 0.5808311487459865; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 22 · 0.5616897313679421; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 24 · 0.47783659480152735; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 25 · 0.2821666630983059; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 31 · 0.49152040367063765; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 33 · 0.5668735513918998; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 36 · 0.5866893299751684; below the displayed range, marked at the axis.Claude Opus 5 · Claude Code · submission 38 · 0.5769691386941643; below the displayed range, marked at the axis.round 139Macro action scoreClaude Opus 5 · Claude Code
Claude Opus 5 · 39 submissions · best 0.6057 · ▽ below 0.592
Claude Opus 5 — every submission, in order (39)
#Score
10.5966
20.5983
30.5624
40.3101
50.5932
60.5889
70.5965
80.6057
90.5997
100.5934
110.6036
120.6040
130.6033
140.5973
150.6057
160.6013
170.5716
180.5721
190.6057
200.6057
210.5808
220.5617
230.6057
240.4778
250.2822
260.6057
270.5959
280.6057
290.5993
300.6057
310.4915
320.6057
330.5669
340.6057
350.6044
360.5867
370.6057
380.5770
390.6057
0.5920.5960.60.6040.6080.6053GPT-5.6 Sol · Codex · submission 3 · 0.5601990639390262; below the displayed range, marked at the axis.GPT-5.6 Sol · Codex · submission 12 · 0.5905210272412961; below the displayed range, marked at the axis.GPT-5.6 Sol · Codex · submission 17 · 0.5735986250297; below the displayed range, marked at the axis.GPT-5.6 Sol · Codex · submission 18 · 0.5533460410634136; below the displayed range, marked at the axis.GPT-5.6 Sol · Codex · submission 19 · 0.3039013154057625; below the displayed range, marked at the axis.GPT-5.6 Sol · Codex · submission 22 · 0.48974277196778726; below the displayed range, marked at the axis.GPT-5.6 Sol · Codex · submission 23 · 0.5574342740981484; below the displayed range, marked at the axis.GPT-5.6 Sol · Codex · submission 24 · 0.5887556372220799; below the displayed range, marked at the axis.GPT-5.6 Sol · Codex · submission 28 · 0.5740402176998849; below the displayed range, marked at the axis.round 140Macro action scoreGPT-5.6 Sol · Codex
GPT-5.6 Sol · 40 submissions · best 0.6053 · ▽ below 0.592
GPT-5.6 Sol — every submission, in order (40)
#Score
10.5966
20.5985
30.5602
40.5940
50.5985
60.5975
70.5991
80.6011
90.5959
100.6001
110.6011
120.5905
130.6009
140.6011
150.6020
160.6022
170.5736
180.5533
190.3039
200.6020
210.6031
220.4897
230.5574
240.5888
250.6044
260.6028
270.6044
280.5740
290.6040
300.6049
310.6031
320.6049
330.5971
340.6043
350.5993
360.5995
370.6053
380.5988
390.5985
400.6039

Links