This page applies the curriculum, day by day, to one real kernel: the fused MLP from the V-JEPA 2 NKI kernels (mlp.py). Every number was derived by hand and then checked against a hardware profile. Use it as a template: swap in another op's shapes and code, and every step carries over.

The workload

V-JEPA 2 ViT-L, one MLP block: out = gelu(x W₁ + b₁) W₂ + b₂, with D = 1024 and F = 4096. One forward is 16 frames at 256×256 at batch 1, cut into 16×16 patches and tubelets 2 frames deep:

R = (256/16)² patches × (16/2) tubelets = 256 × 8 = 2048 tokens

Two versions are compared: an unfused torch.compile MLP, and the fused NKI kernel mlp.py, which keeps the GELU output in SBUF and runs on one logical core (LNC=2).

Results: hand count against hardware

Hand count Measured
Fused mlp.py: FLOPs 34.360 G 34.360 G
Fused mlp.py: intensity 819.2 789
torch.compile MLP: FLOPs 34.360 G 39.729 G
torch.compile MLP: intensity 315.1 314.8

Day 1: The roofline

W = 34,359,738,368 FLOPs     P  = 157.3e12 FLOP/s
Q = 41,943,040 bytes         BW = 0.75e12 bytes/s

T_compute = W ÷ P  = 218.5 µs
T_memory  = Q ÷ BW =  55.9 µs      →  218.5 µs ≤ T ≤ 274.4 µs
I         = W ÷ Q  = 819.2 FLOPs/byte
I × BW    = 614 TFLOP/s  >  P = 157.3   →  compute-bound

The unfused version: I = 315.1, so I × BW = 236 TFLOP/s, also above P. Both are compute-bound; the ridge is P ÷ BW ≈ 210.

Day 2: FLOPs and bytes from the shapes

The MLP is two matmuls, [R × D] by [D × F] and [R × F] by [F × D]:

W = 2RDF + 2RFD = 4RDF = 4 × 2048 × 1024 × 4096 = 34,359,738,368 FLOPs

The fused kernel splits the rows across 2 cores, so each handles r = R ÷ 2 = 1024 rows and loads both weight matrices (Day 6). Per core:

$$ I = \frac{4rDF}{p\,(2DF + 2rD)} = \frac{rF}{F + r} = \frac{1024 \times 4096}{5120} = 819.2 $$

With weights alone it would be exactly r = 1024: each weight is reused once per row. The shape matters: at R = 256 (r = 128) the same kernel gives I = 124, below the ridge, so memory-bound. Setting rF ÷ (F + r) = 210 gives r ≈ 221, so the kernel crosses the ridge at about 442 tokens.

Day 3: Peaks from the hardware

P per core       = 128 × 128 × 2 × 2.4e9 = 78.6e12 FLOP/s   (documented: 79 TFLOPS)
P for LNC=2      = 2 × 78.6               = 157.3e12 FLOP/s  (the kernel card's MFU denominator)
BW per core pair = 2 × 3e12 ÷ 8 cores    = 0.75e12 bytes/s
Ridge            = 157.3 ÷ 0.75           ≈ 210 FLOPs/byte

Day 4: Measure, and compare the parts