All logged agent runs share one plot. Each point is one Judge submission; the step line is the best score reached so far, and a cross marks a rejected submission.
View by
The agent runs · 2 runs · submissionrunning bestbaseline · Judge measuredTime since run start · points mark recorded Judge results
Find an inference strategy that helps the fixed Molmo2-VideoPoint model locate visual anomalies more accurately in generated videos. The model is not trained or fine-tuned; the experiment changes only how inference is run.
Environment
The model, videos, questions, and evaluator are already provided.
The agent delivers one small strategy file describing its inference choices.
Reference baseline
Every submission is evaluated on the same reserved videos and receives an absolute soft-F1 score. The shipped strategy provides a common starting point; the Judge does not rerun a separate baseline beside every candidate.
Research loop
Inspect the current strategy and the feedback from earlier submissions.
Change one inference choice for a clear reason, validate it, and submit it.
Compare the aggregate result, then keep, revise, or reject the idea.
What the agent may change
How frames are selected across the video.
How densely the video is sampled.
How many frames the model receives.
How the pointing request is phrased.
How long the model may continue its answer.
What stays fixed
The Molmo2 model and all of its weights.
The videos, questions, labels, and reserved evaluation cases.
The coordinate parser, output rules, metric, and scoring procedure.
Deterministic generation and the isolated single-GPU runtime.
The declarative strategy-file interface; candidates cannot add executable code.
Evaluation
The Judge reports mean spatiotemporal soft-F1 on the fixed reserved set; higher is better. The agent sees aggregate quality and validity signals after each submission, while per-video targets and predictions remain hidden.
Agent runs
Each figure is one run: a lollipop per submission, the running best as a step line, and a cross on the floor where a submission was rejected. Open the row below each figure for every Judge submission.
Claude Fable 5 · 60 submissions · best 0.462Claude Fable 5 — every submission, in order (60)
#
Score
1
0.32
2
0.358
3
0.334
4
0.358
5
0.358
6
0.338
7
0.358
8
0.358
9
0.348
10
0.435
11
0.426
12
0.359
13
0.35
14
0.427
15
0.378
16
0.378
17
0.431
18
0.444
19
0.41
20
0.367
21
0.43
22
0.438
23
0.444
24
0.444
25
0.444
26
0.427
27
0.309
28
0.309
29
0.411
30
0.444
31
0.378
32
0.384
33
0.424
34
0.403
35
0.444
36
0.431
37
0.427
38
0.426
39
0.444
40
0.349
41
0.342
42
0.448
43
0.431
44
0.402
45
0.444
46
0.392
47
0.431
48
0.419
49
0.462
50
0.448
51
0.434
52
0.434
53
0.419
54
0.414
55
0.406
56
0.462
57
0.376
58
0.405
59
0.362
60
—
GPT-5.6 Sol · 56 submissions · best 0.462GPT-5.6 Sol — every submission, in order (56)