This page applies the curriculum, day by day, to one real kernel: the fused MLP from the V-JEPA 2 NKI kernels (mlp.py). Every number was derived by hand and then checked against a hardware profile. Use it as a template: swap in another op's shapes and code, and every step carries over.
V-JEPA 2 ViT-L, one MLP block: out = gelu(x W₁ + b₁) W₂ + b₂, with D = 1024 and F = 4096. One forward is 16 frames at 256×256 at batch 1, cut into 16×16 patches and tubelets 2 frames deep:
R = (256/16)² patches × (16/2) tubelets = 256 × 8 = 2048 tokens
Two versions are compared: an unfused torch.compile MLP, and the fused NKI kernel mlp.py, which keeps the GELU output in SBUF and runs on one logical core (LNC=2).
| Hand count | Measured | |
|---|---|---|
Fused mlp.py: FLOPs |
34.360 G | 34.360 G |
Fused mlp.py: intensity |
819.2 | 789 |
| torch.compile MLP: FLOPs | 34.360 G | 39.729 G |
| torch.compile MLP: intensity | 315.1 | 314.8 |
W = 34,359,738,368 FLOPs P = 157.3e12 FLOP/s
Q = 41,943,040 bytes BW = 0.75e12 bytes/s
T_compute = W ÷ P = 218.5 µs
T_memory = Q ÷ BW = 55.9 µs → 218.5 µs ≤ T ≤ 274.4 µs
I = W ÷ Q = 819.2 FLOPs/byte
I × BW = 614 TFLOP/s > P = 157.3 → compute-bound
The unfused version: I = 315.1, so I × BW = 236 TFLOP/s, also above P. Both are compute-bound; the ridge is P ÷ BW ≈ 210.
The MLP is two matmuls, [R × D] by [D × F] and [R × F] by [F × D]:
W = 2RDF + 2RFD = 4RDF = 4 × 2048 × 1024 × 4096 = 34,359,738,368 FLOPs
The fused kernel splits the rows across 2 cores, so each handles r = R ÷ 2 = 1024 rows and loads both weight matrices (Day 6). Per core:
$$ I = \frac{4rDF}{p\,(2DF + 2rD)} = \frac{rF}{F + r} = \frac{1024 \times 4096}{5120} = 819.2 $$
With weights alone it would be exactly r = 1024: each weight is reused once per row. The shape matters: at R = 256 (r = 128) the same kernel gives I = 124, below the ridge, so memory-bound. Setting rF ÷ (F + r) = 210 gives r ≈ 221, so the kernel crosses the ridge at about 442 tokens.
P per core = 128 × 128 × 2 × 2.4e9 = 78.6e12 FLOP/s (documented: 79 TFLOPS)
P for LNC=2 = 2 × 78.6 = 157.3e12 FLOP/s (the kernel card's MFU denominator)
BW per core pair = 2 × 3e12 ÷ 8 cores = 0.75e12 bytes/s
Ridge = 157.3 ÷ 0.75 ≈ 210 FLOPs/byte