Task Sample · Systems

MInference 32-Head Sparse Prefill Kernel

Speed up a fixed sparse-attention prefill operator without changing which tokens it attends to or weakening its numerical result.

Best score per run18-case geometric-mean speedup (fallback median ms / candidate median ms)
01.66
Model · harnessBestSubmissionsRuntime
GPT-5.6 SolCodex · high1.54616.31 h
Claude Opus 5Claude Code · high1.5156.02 h

Sparse-prefill latency is measured against the frozen fallback inside the same Judge run.

Teaser figure from MInference
Source project · MInference · source ↗

The climb

All logged agent runs share one plot. Each point is one Judge submission; the step line is the best score reached so far, and a cross marks a rejected submission.

0.8811.081.271.461.66Baseline · 1.000×0246.4Elapsed time (h)18-case geometric-mean speedupGPT-5.6 Sol · submission 1 · 1.08 · 0.167 h, Judge result recordedGPT-5.6 Sol · submission 2 · 0 · 0.267 h, Judge result recordedGPT-5.6 Sol · submission 3 · 1.14 · 0.318 h, Judge result recordedGPT-5.6 Sol · submission 4 · 1.14 · 0.385 h, Judge result recordedGPT-5.6 Sol · submission 5 · 1.19 · 0.455 h, Judge result recordedGPT-5.6 Sol · submission 6 · 1.2 · 0.569 h, Judge result recordedGPT-5.6 Sol · submission 7 · 1.24 · 0.884 h, Judge result recordedGPT-5.6 Sol · submission 8 · 1.31 · 0.935 h, Judge result recordedGPT-5.6 Sol · submission 9 · 1.33 · 1.15 h, Judge result recordedGPT-5.6 Sol · submission 10 · 1.33 · 1.25 h, Judge result recordedGPT-5.6 Sol · submission 11 · 1.34 · 1.34 h, Judge result recordedGPT-5.6 Sol · submission 12 · 1.36 · 1.41 h, Judge result recordedGPT-5.6 Sol · submission 13 · 1.35 · 1.52 h, Judge result recordedGPT-5.6 Sol · submission 14 · 1.36 · 1.59 h, Judge result recordedGPT-5.6 Sol · submission 15 · 1.36 · 1.66 h, Judge result recordedGPT-5.6 Sol · submission 16 · 1.36 · 1.71 h, Judge result recordedGPT-5.6 Sol · submission 17 · 1.39 · 1.79 h, Judge result recordedGPT-5.6 Sol · submission 18 · 1.39 · 1.9 h, Judge result recordedGPT-5.6 Sol · submission 19 · 1.39 · 1.99 h, Judge result recordedGPT-5.6 Sol · submission 20 · 1.4 · 2.06 h, Judge result recordedGPT-5.6 Sol · submission 21 · 1.39 · 2.11 h, Judge result recordedGPT-5.6 Sol · submission 22 · 1.38 · 2.17 h, Judge result recordedGPT-5.6 Sol · submission 23 · 1.46 · 2.23 h, Judge result recordedGPT-5.6 Sol · submission 24 · 1.46 · 2.28 h, Judge result recordedGPT-5.6 Sol · submission 25 · 1.46 · 2.38 h, Judge result recordedGPT-5.6 Sol · submission 26 · 1.47 · 2.5 h, Judge result recordedGPT-5.6 Sol · submission 27 · 1.45 · 2.54 h, Judge result recordedGPT-5.6 Sol · submission 28 · 1.47 · 2.62 h, Judge result recordedGPT-5.6 Sol · submission 29 · 1.46 · 2.68 h, Judge result recordedGPT-5.6 Sol · submission 30 · 0 · 3.04 h, Judge result recordedGPT-5.6 Sol · submission 31 · 1.47 · 3.08 h, Judge result recordedGPT-5.6 Sol · submission 32 · 1.46 · 3.17 h, Judge result recordedGPT-5.6 Sol · submission 33 · 1.46 · 3.22 h, Judge result recordedGPT-5.6 Sol · submission 34 · 1.46 · 3.3 h, Judge result recordedGPT-5.6 Sol · submission 35 · 1.46 · 3.4 h, Judge result recordedGPT-5.6 Sol · submission 36 · 1.47 · 3.47 h, Judge result recordedGPT-5.6 Sol · submission 37 · 1.47 · 3.56 h, Judge result recordedGPT-5.6 Sol · submission 38 · 1.47 · 3.61 h, Judge result recordedGPT-5.6 Sol · submission 39 · 1.48 · 3.72 h, Judge result recordedGPT-5.6 Sol · submission 40 · 1.47 · 3.82 h, Judge result recordedGPT-5.6 Sol · submission 41 · 1.48 · 3.87 h, Judge result recordedGPT-5.6 Sol · submission 42 · 1.47 · 3.93 h, Judge result recordedGPT-5.6 Sol · submission 43 · 1.48 · 4 h, Judge result recordedGPT-5.6 Sol · submission 44 · 1.48 · 4.04 h, Judge result recordedGPT-5.6 Sol · submission 45 · 1.47 · 4.1 h, Judge result recordedGPT-5.6 Sol · submission 46 · 1.48 · 4.18 h, Judge result recordedGPT-5.6 Sol · submission 47 · 1.48 · 4.23 h, Judge result recordedGPT-5.6 Sol · submission 48 · 1.48 · 4.29 h, Judge result recordedGPT-5.6 Sol · submission 49 · 1.47 · 4.36 h, Judge result recordedGPT-5.6 Sol · submission 50 · 1.47 · 4.43 h, Judge result recordedGPT-5.6 Sol · submission 51 · 1.47 · 4.72 h, Judge result recordedGPT-5.6 Sol · submission 52 · 1.53 · 4.91 h, Judge result recordedGPT-5.6 Sol · submission 53 · 1.54 · 5.04 h, Judge result recordedGPT-5.6 Sol · submission 54 · 1.54 · 5.16 h, Judge result recordedGPT-5.6 Sol · submission 55 · 1.53 · 5.54 h, Judge result recordedGPT-5.6 Sol · submission 56 · 1.53 · 5.61 h, Judge result recordedGPT-5.6 Sol · submission 57 · 1.54 · 5.69 h, Judge result recordedGPT-5.6 Sol · submission 58 · 1.53 · 5.8 h, Judge result recordedGPT-5.6 Sol · submission 59 · 1.54 · 5.95 h, Judge result recordedGPT-5.6 Sol · submission 60 · 1.54 · 6.26 h, Judge result recordedGPT-5.6 Sol · submission 61 · 1.54 · 6.31 h, Judge result recordedGPT-5.6 Sol · 1.54Claude Opus 5 · submission 1 · 1.23 · 0.483 h, Judge result recordedClaude Opus 5 · submission 2 · 0 · 1.12 h, Judge result recordedClaude Opus 5 · submission 3 · 1.37 · 1.22 h, Judge result recordedClaude Opus 5 · submission 4 · 1.45 · 1.34 h, Judge result recordedClaude Opus 5 · submission 5 · 1.47 · 3.58 h, Judge result recordedClaude Opus 5 · submission 6 · 1.46 · 3.77 h, Judge result recordedClaude Opus 5 · submission 7 · 1.47 · 3.92 h, Judge result recordedClaude Opus 5 · submission 8 · 1.5 · 4.1 h, Judge result recordedClaude Opus 5 · submission 9 · 1.49 · 4.25 h, Judge result recordedClaude Opus 5 · submission 10 · 1.49 · 4.31 h, Judge result recordedClaude Opus 5 · submission 11 · 1.49 · 4.54 h, Judge result recordedClaude Opus 5 · submission 12 · 1.5 · 4.7 h, Judge result recordedClaude Opus 5 · submission 13 · 1.5 · 4.94 h, Judge result recordedClaude Opus 5 · submission 14 · 1.5 · 5.56 h, Judge result recordedClaude Opus 5 · submission 15 · 1.49 · 6.02 h, Judge result recordedClaude Opus 5 · 1.5
The agent runs · 2 runs · submissionrejected submissionrunning bestbaseline · paired normalizedTime since run start · points mark recorded Judge results

The task

Reduce the latency of a fixed vertical-plus-slash sparse-attention prefill operation without dropping attention work, changing causality, or weakening its output.

Environment

Reference baseline

Every submission runs the frozen upstream fallback and the candidate on matching fresh inputs in the same Judge invocation. The score is therefore a paired speedup rather than an absolute latency.

Research loop

  1. Inspect the fallback, sparse-attention contract, public diagnostics, and earlier aggregate feedback.
  2. Test one execution-plan hypothesis and record the public evidence.
  3. Finish the source, submit it, and compare latency, correctness, and resource aggregates before revising.

What the agent may change

What stays fixed

Evaluation

For every reserved case, the Judge compares fallback and candidate median latency and combines the ratios geometrically; higher is better. Scope, stream, determinism, full-write, mutation, and numerical checks must pass before the speed result is accepted.

Agent runs

Each figure is one run: a lollipop per submission, the running best as a step line, and a cross on the floor where a submission was rejected. Open the row below each figure for every Judge submission.

0.9971.151.311.471.621.54round 16118-case geometric-mean speedupGPT-5.6 Sol · Codex
GPT-5.6 Sol · 61 submissions · best 1.54
GPT-5.6 Sol — every submission, in order (61)
#Score
11.08
20
31.14
41.14
51.19
61.2
71.24
81.31
91.33
101.33
111.34
121.36
131.35
141.36
151.36
161.36
171.39
181.39
191.39
201.4
211.39
221.38
231.46
241.46
251.46
261.47
271.45
281.47
291.46
300
311.47
321.46
331.46
341.46
351.46
361.47
371.47
381.47
391.48
401.47
411.48
421.47
431.48
441.48
451.47
461.48
471.48
481.48
491.47
501.47
511.47
521.53
531.54
541.54
551.53
561.53
571.54
581.53
591.54
601.54
611.54
1.181.271.371.461.551.5round 11518-case geometric-mean speedupClaude Opus 5 · Claude Code
Claude Opus 5 · 15 submissions · best 1.5
Claude Opus 5 — every submission, in order (15)
#Score
11.23
20
31.37
41.45
51.47
61.46
71.47
81.5
91.49
101.49
111.49
121.5
131.5
141.5
151.49
GPT-5.6 Sol — the agent’s own closing summary

Implemented and submitted the optimized operator. Best formal result: 1.5405066467× geometric-mean speedup (submission 60).

  • 32K: 1.5770×
  • 65K: 1.5369×
  • 131K: 1.5084×
  • Correctness and policy checks passed
  • Peak HBM: 16,385.25 MiB

The retained plan uses certified translated-quad fusion, excludes boundary-crossing final quads, and removes redundant tail activity predicates through an exact causal-range certificate.

Claude Opus 5 — the agent’s own closing summary

Delivered: /workspace/submission.py at approximately 1.50× the frozen fallback, with best recorded reward 1.5041 (32K: 1.519×, 64K: 1.524×, 128K: 1.461×), plus a summary-first notes.md covering 23 hypotheses — 5 shipped and 18 rejected on measurement.

The retained gains came from causal-mask hoisting, a four-way merged path, closed-form index conversion where valid, a streaming converter elsewhere, and heaviest-first program ordering.

Links