Task Sample · Systems

H100 fp16 GEMM Kernel Lab

Build and refine a clean-room CUDA matrix-multiplication kernel for higher sustained H100 throughput while preserving the fixed interface and correctness rules.

Best score per runMedian TFLOP/s over five fixed-protocol repeats (4096x4096x4096 fp16 GEMM)
0843
Model · harnessBestSubmissionsRuntime
Claude Opus 5Claude Code · high780375.33 h
GPT-5.6 SolCodex · high619346.07 h

Absolute median valid throughput on the reserved H100 workload; higher is better.

Teaser figure from OpenRSI Index task `task_collect/zy_tasks/gemm-h100-kernel-lab` (EdgeBench semantic source `tasks/archive/gemm-h100-edgebench`)
Source project · OpenRSI Index task `task_collect/zy_tasks/gemm-h100-kernel-lab` · source ↗

The climb

All logged agent runs share one plot. Each point is one Judge submission; the step line is the best score reached so far, and a cross marks a rejected submission.

02505007501000Baseline · 196.8 TFLOP/s0246.1Elapsed time (h)Median TFLOP/s over five…Claude Opus 5 · submission 1 · 499 · 0.3 h, Judge result recordedClaude Opus 5 · submission 2 · 661 · 0.371 h, Judge result recordedClaude Opus 5 · submission 3 · 727 · 0.504 h, Judge result recordedClaude Opus 5 · submission 4 · 739 · 0.953 h, Judge result recordedClaude Opus 5 · submission 5 · 739 · 1.13 h, Judge result recordedClaude Opus 5 · submission 6 · 742 · 1.17 h, Judge result recordedClaude Opus 5 · submission 7 · 740 · 1.29 h, Judge result recordedClaude Opus 5 · submission 8 · 749 · 3.44 h, Judge result recordedClaude Opus 5 · submission 9 · 749 · 3.49 h, Judge result recordedClaude Opus 5 · submission 10 · 749 · 3.58 h, Judge result recordedClaude Opus 5 · submission 11 · 751 · 3.6 h, Judge result recordedClaude Opus 5 · submission 12 · 751 · 3.65 h, Judge result recordedClaude Opus 5 · submission 13 · 753 · 3.7 h, Judge result recordedClaude Opus 5 · submission 14 · 749 · 3.79 h, Judge result recordedClaude Opus 5 · submission 15 · 749 · 4 h, Judge result recordedClaude Opus 5 · submission 16 · 750 · 4.05 h, Judge result recordedClaude Opus 5 · submission 17 · 749 · 4.12 h, Judge result recordedClaude Opus 5 · submission 18 · 750 · 4.19 h, Judge result recordedClaude Opus 5 · submission 19 · 749 · 4.26 h, Judge result recordedClaude Opus 5 · submission 20 · 749 · 4.31 h, Judge result recordedClaude Opus 5 · submission 21 · 748 · 4.34 h, Judge result recordedClaude Opus 5 · submission 22 · 777 · 4.53 h, Judge result recordedClaude Opus 5 · submission 23 · 780 · 4.64 h, Judge result recordedClaude Opus 5 · submission 24 · 779 · 4.75 h, Judge result recordedClaude Opus 5 · submission 25 · 775 · 4.78 h, Judge result recordedClaude Opus 5 · submission 26 · 775 · 4.82 h, Judge result recordedClaude Opus 5 · submission 27 · 776 · 4.86 h, Judge result recordedClaude Opus 5 · submission 28 · 775 · 4.9 h, Judge result recordedClaude Opus 5 · submission 29 · 775 · 4.93 h, Judge result recordedClaude Opus 5 · submission 30 · 774 · 4.98 h, Judge result recordedClaude Opus 5 · submission 31 · 774 · 5.01 h, Judge result recordedClaude Opus 5 · submission 32 · 771 · 5.03 h, Judge result recordedClaude Opus 5 · submission 33 · 777 · 5.1 h, Judge result recordedClaude Opus 5 · submission 34 · 778 · 5.16 h, Judge result recordedClaude Opus 5 · submission 35 · 771 · 5.28 h, Judge result recordedClaude Opus 5 · submission 36 · 777 · 5.31 h, Judge result recordedClaude Opus 5 · submission 37 · 777 · 5.33 h, Judge result recordedClaude Opus 5 · 780GPT-5.6 Sol · submission 1 · 257 · 0.147 h, Judge result recordedGPT-5.6 Sol · submission 2 · 275 · 0.323 h, Judge result recordedGPT-5.6 Sol · submission 3 · 295 · 0.354 h, Judge result recordedGPT-5.6 Sol · submission 4 · 297 · 0.418 h, Judge result recordedGPT-5.6 Sol · submission 5 · 308 · 0.79 h, Judge result recordedGPT-5.6 Sol · submission 6 · 394 · 1.28 h, Judge result recordedGPT-5.6 Sol · submission 7 · 399 · 1.34 h, Judge result recordedGPT-5.6 Sol · submission 8 · 436 · 1.98 h, Judge result recordedGPT-5.6 Sol · submission 9 · 519 · 2.26 h, Judge result recordedGPT-5.6 Sol · submission 10 · 520 · 2.4 h, Judge result recordedGPT-5.6 Sol · submission 11 · 531 · 2.7 h, Judge result recordedGPT-5.6 Sol · submission 12 · 533 · 2.77 h, Judge result recordedGPT-5.6 Sol · submission 13 · 536 · 2.82 h, Judge result recordedGPT-5.6 Sol · submission 14 · 567 · 2.88 h, Judge result recordedGPT-5.6 Sol · submission 15 · 572 · 3.27 h, Judge result recordedGPT-5.6 Sol · submission 16 · 571 · 3.4 h, Judge result recordedGPT-5.6 Sol · submission 17 · 575 · 3.43 h, Judge result recordedGPT-5.6 Sol · submission 18 · 585 · 3.57 h, Judge result recordedGPT-5.6 Sol · submission 19 · 590 · 3.64 h, Judge result recordedGPT-5.6 Sol · submission 20 · 588 · 3.73 h, Judge result recordedGPT-5.6 Sol · submission 21 · 590 · 3.78 h, Judge result recordedGPT-5.6 Sol · submission 22 · 593 · 4.31 h, Judge result recordedGPT-5.6 Sol · submission 23 · 594 · 4.38 h, Judge result recordedGPT-5.6 Sol · submission 24 · 597 · 4.62 h, Judge result recordedGPT-5.6 Sol · submission 25 · 597 · 4.81 h, Judge result recordedGPT-5.6 Sol · submission 26 · 596 · 4.85 h, Judge result recordedGPT-5.6 Sol · submission 27 · 601 · 4.91 h, Judge result recordedGPT-5.6 Sol · submission 28 · 601 · 5.02 h, Judge result recordedGPT-5.6 Sol · submission 29 · 601 · 5.12 h, Judge result recordedGPT-5.6 Sol · submission 30 · 600 · 5.26 h, Judge result recordedGPT-5.6 Sol · submission 31 · 616 · 5.6 h, Judge result recordedGPT-5.6 Sol · submission 32 · 617 · 5.9 h, Judge result recordedGPT-5.6 Sol · submission 33 · 616 · 6.03 h, Judge result recordedGPT-5.6 Sol · submission 34 · 619 · 6.07 h, Judge result recordedGPT-5.6 Sol · 619
The agent runs · 2 runs · submissionrunning bestbaseline · local reference runTime since run start · points mark recorded Judge results

The task

Write a clean-room CUDA kernel that multiplies row-major half-precision matrices correctly and sustains as much throughput as possible on the fixed H100 workload.

Environment

Reference baseline

A simple clean-room tensor-core kernel is the starting point. It is not measured beside every candidate, so the published run scores are absolute throughput results.

Research loop

  1. Inspect the current kernel family and earlier measurements.
  2. Test one kernel or scheduling hypothesis in a new version.
  3. Build, check correctness, measure, submit the selected version, and use the result to choose the next idea.

What the agent may change

What stays fixed

Evaluation

The Judge builds the selected source, checks that it computes the full result correctly, and reports median valid throughput; higher is better. Candidate-owned build or correctness failures do not receive a normal performance result.

Agent runs

Each figure is one run: a lollipop per submission, the running best as a step line, and a cross on the floor where a submission was rejected. Open the row below each figure for every Judge submission.

449544640735831780round 137Median TFLOP/s over five…Claude Opus 5 · Claude Code
Claude Opus 5 · 37 submissions · best 780
Claude Opus 5 — every submission, in order (37)
#Score
1499
2661
3727
4739
5739
6742
7740
8749
9749
10749
11751
12751
13753
14749
15749
16750
17749
18750
19749
20749
21748
22777
23780
24779
25775
26775
27776
28775
29775
30774
31774
32771
33777
34778
35771
36777
37777
192315438561684619round 134Median TFLOP/s over five…GPT-5.6 Sol · Codex
GPT-5.6 Sol · 34 submissions · best 619
GPT-5.6 Sol — every submission, in order (34)
#Score
1257
2275
3295
4297
5308
6394
7399
8436
9519
10520
11531
12533
13536
14567
15572
16571
17575
18585
19590
20588
21590
22593
23594
24597
25597
26596
27601
28601
29601
30600
31616
32617
33616
34619
Claude Opus 5 — the agent’s own closing summary

Final deliverable: FINAL_VERSION = 016, best score 780.2 TFLOP/s (4.0× the WMMA baseline), 17 candidates, a 957-line notes.md, a clean workspace, and idle GPUs.

I resolved 22 hypotheses — 8 shipped and 14 refuted — with each remaining path either below the protocol's noise floor or blocked by a measured register, shared-memory, bandwidth, or launch-overhead limit.

GPT-5.6 Sol — the agent’s own closing summary

Selected v193 as the final kernel.

  • Reserved Judge result: 616.517 TFLOP/s, five valid repeats
  • Improvement over v164: +15.369 TFLOP/s
  • Correctness, full overwrite, fallback shapes, stream behavior, and source admission all pass
  • Final exact local verification: 609.950 TFLOP/s, zero bad or unwritten samples

The main breakthrough was passing immutable TMA descriptors through __grid_constant__ kernel parameters, eliminating per-CTA descriptor publication. The final version also omits the unnecessary explicit wait on the terminal committed TMA output store.

Links