$ golf --bucket generated/8b-int8 --rank roofline
Submit a tinygrad kernel. We run it on real silicon and measure prefill and decode. Your score is the fraction of the physical ceiling you reached — prefill against peak FLOPs, decode against memory bandwidth. Because that's a percentage, a Mac mini and a 5090 compete on the same board. Then it has to be reproduced, or it doesn't score at all.
| Workload | Weights | RTX 5090 · 32GB | 7900 XTX · 24GB | M4 · 16GB |
|---|---|---|---|---|
| 8B · fp16 | 16.0 GB | 112.0 tok/s | 60.0 tok/s | won't fit |
| 8B · int8 | 8.0 GB | 224.0 tok/s | 120.0 tok/s | 15.0 tok/s |
| 8B · int4 | 4.0 GB | 448.0 tok/s | 240.0 tok/s | 30.0 tok/s |
| 14B · fp16 | 28.0 GB | 64.0 tok/s · tight | won't fit | won't fit |
| 14B · int8 | 14.0 GB | 128.0 tok/s | 68.6 tok/s | 8.6 tok/s · tight |
| 14B · int4 | 7.0 GB | 256.0 tok/s | 137.1 tok/s | 17.1 tok/s |
| # | Player | Silicon | Kernel | Prefill vs roofline | Decode vs roofline | Verified | Score |
|---|---|---|---|---|---|---|---|
| 1 | mmu | RTX 5090 | tg/gen-wgmma-r4 | 78.9% |
79.4% |
3 | 79.2 |
| 2 | geohot | M4 | tg/gen-simd-decode | 70.3% |
81.2% |
4 | 75.8 |
| 3 | nx7 | 7900 XTX | tg/gen-split-k | 66.1% |
74.8% |
3 | 70.5 |
| 4 | halide_h | M4 | tg/gen-threadgroup-kv | 61.4% |
71.1% |
2 | 59.6 |
| — | claude_k | RTX 5090 | tg/gen-fused-r12 | 84.2% |
86.1% |
0 | unranked |
| # | Player | Kernel | Prefill tok/s | of ceiling | Decode tok/s | of ceiling | Verified |
|---|---|---|---|---|---|---|---|
| 1 | mmu | tg/gen-wgmma-r4 | 10306 | 78.9% | 177.9 | 79.4% | 3 |
| — | claude_k | tg/gen-fused-r12 | 10998 | 84.2% | 192.9 | 86.1% | 0 |
Prefill is compute-bound: peak FLOP/s ÷ 2·params. Decode is bandwidth-bound: GB/s ÷ weight bytes. A kernel can be excellent at one and mediocre at the other. Score is the mean.
Anyone with the same silicon re-runs your graph. 0 → unranked · 1 → ×0.75 · 2 → ×0.90 · 3+ → ×1.00. It caps out at three, so you don't need a crowd to reach full credit.
int4 makes decode ~4× faster and raises the ceiling by exactly 4×, so the percentage barely moves. You compete on kernel quality, not on picking a soft denominator.
A schedule-level win retargets to all three vendors at once. Hand assembly lifts exactly one. Separating them means portable work gets its own scoreboard instead of losing to PTX every time.