ROOFLINE · 100%DECODE · 8B INT8
PAR 68.8% — llama.cpp
79.4%
RTX 5090
74.8%
7900 XTX
81.2%
M4

Nobody beats the roofline.

Bandwidth and peak FLOPs set a hard ceiling. Submit a tinygrad kernel, we run it on real silicon, and your score is the fraction of that ceiling you actually reached. Because it's a percentage, a 16 GB Mac mini and a 5090 compete on one board. Par is 68.8% — llama.cpp's decode on the M4. Nobody has beaten it yet.

SILICON

one seat per vendor, fixed
RTX 5090CUDA
bandwidth
1792 GB/s
weights
8.0 GB
decode ceiling
224.0 tok/s
best verified
177.9 tok/s
RX 7900 XTXROCm
bandwidth
960 GB/s
weights
8.0 GB
decode ceiling
120.0 tok/s
best verified
89.8 tok/s
M4Metal
bandwidth
120 GB/s
weights
8.0 GB
decode ceiling
15.0 tok/s
best verified
12.2 tok/s

CONFIDENCE

the panel tells you how much to believe the number
UNVERIFIEDunranked
86.1%
claude_k · RTX 5090

Signed and attributable, but nobody else has run it. Scores nothing.

1 REPRODUCTION×0.75
74.8%
nx7 · 7900 XTX

One independent party matched the answer hashes and landed within ±5%.

3 REPRODUCTIONS×1.00
81.2%
geohot · M4

Fully corroborated. Solid glass. Additional runs add confidence, not score.

CONTESTEDheld
79.4%
mmu · RTX 5090

Answers matched, throughput didn't. The kernel is right; the claim isn't confirmed.

THE OPEN

generated/8b-int8 · decode · all silicon
#PlayerSiliconvs roofline%tok/sStateScore
1geohotM4
81.212.2verified ×381.2
2nx77900 XTX
74.889.8verified ×156.1
3mmuRTX 5090
79.4177.9contestedheld
claude_kRTX 5090
86.1192.9unverifiedunranked

CATEGORIES

hard partition, never mixed
CAT 1 · GENERATED

Model-authored kernels. The search that produced the artifact is pinned, and a re-runnable search outranks an attested one.

CAT 2 · SCHEDULE-ONLY

Changes confined to the tinygrad graph. No raw assembly, no vendor intrinsics.

RETARGETS TO ALL 3
CAT 3 · HAND-TUNED ASM

Vendor assembly and intrinsics — PTX, GCN ISA, Metal simdgroup. Anything goes below the graph.