This guide is the whole process in one pass, in the order you would actually do it: start with a chip and an op, end with a plotted point and an explanation of why it sits where it does. The lessons page explains why each idea is true. This page explains what to do, what to write down, and how to tell when a step has gone wrong.

One example runs through every step so the numbers connect: a bf16 matmul $C = A\,B$ with $M = K = N = 4096$ on one Trainium2 NeuronCore. Every figure for it below was computed by the script in Step 10.

Notation

Symbol Meaning Units
$P$ Peak compute: the most arithmetic the hardware can perform per second, for one dtype FLOP/s
$BW$ Bandwidth: the most data the slow link, here HBM into SBUF, can move per second bytes/s
$W$ Work: the arithmetic the kernel performs FLOPs
$I$ Arithmetic intensity: FLOPs divided by the bytes moved across the slow link, counting every reload FLOPs per byte
$I^{*}$ Ridge point, P ÷ BW: the intensity where the two roofs meet FLOPs per byte
$M,\ K,\ N$ Shape of the matmul C = A · B, with A of shape M × K, B of shape K × N and C of shape M × N. K is the dimension that is summed over. count
$t_{\text{measured}}$ The kernel's execution time on the device, from the profiler seconds

What you need before starting

Step 1: Pin down exactly what is being entitled as

Entitlement is always for one specific thing, and most wrong answers come from leaving one of these four choices vague.

One op with fixed shapes. Write the shapes down as numbers. Here: 4096 × 4096 times 4096 × 4096.

One dtype. Peak compute and bytes per element both depend on it, and they must use the same one. Here: bf16 in and out, 2 bytes per element.

One compute unit. Decide whether the answer is for one physical core, one logical core, or the whole device, and keep the peak, the bandwidth and the byte count at that same scope. A kernel launched as kernel[2] under LNC=2 runs on two physical cores, each with its own SBUF, so its scope is the pair: twice the peak, twice the per-core bandwidth (750 GB/s), and every per-core load counted twice. Here: one physical NeuronCore.

One link. The roofline counts bytes crossing a single slow link. Name it. Here: HBM to SBUF, the path the DMA engines serve. Data already sitting in SBUF is free in this model.

Write these four lines at the top of the page. Every later number should be checkable against them.

Step 2: Derive peak compute from the hardware

Peak compute is a count of arithmetic units multiplied by how often each can fire. Find three things in the architecture document: how many multiply-accumulates the array performs per clock cycle, how many FLOPs each one counts as (a multiply plus an add is 2), and the clock frequency.

For the Trainium2 Tensor Engine, the matmul tile limits are a 128 × 128 stationary tile against a moving tile of 128 × 512. The architecture guide describes the engine as a 128 × 128 systolic array that streams one column of the moving tile per cycle, so it performs 128 × 128 multiply-accumulates per cycle. At the documented 2.4 GHz clock:

$$ P = 128\times128\times2\times2.4\times10^{9} = 7.86\times10^{13}\ \text{FLOP/s} $$