Field note 18 · Jul 29, 2026 · architect

Prefill and decode are two different workloads

  • disaggregated-prefill-decode
  • helios-rackscale
  • wafer-scale-decode-offload
  • vendor-modelled-benchmark
split panel, one pool for both stages vs prefill on Helios / decode on the Wafer-Scale Engine, with both headline numbers and their actual baselines
split panel, one pool for both stages vs prefill on Helios / decode on the Wafer-Scale Engine, with both headline numbers and their actual baselines

Prefill and decode are two different workloads. On July 23 the hardware stopped pretending otherwise.

AMD launched Helios and, the same day, announced a Cerebras pairing that splits inference across two machines: Helios processes the prompts and large context windows, the Wafer-Scale Engine handles token generation. AMD's own term for it is disaggregated inference.

If you have sized an on-prem cluster you already knew the split. Prefill is compute bound and batches well. Decode is memory bandwidth bound and does not, which is why a single pool sized for both means you overpay on prefill and still miss your latency SLO on decode.

Two numbers came out that day. Keep them apart.

30% more tokens per dollar: Helios against NVIDIA Vera Rubin NVL72. AMD Performance Labs estimate, Kimi K2 Thinking, 32K in and 8K out.

5x tokens per second per watt: AMD plus Cerebras against a Cerebras WSE-only configuration. Different model, Kimi 2.6 1T. Vendor modelling, not measured silicon.

The trap: that 5x is not a win over NVIDIA. It is Cerebras with Helios against Cerebras without it, and the headline says per watt while the footnote says per kilowatt.

The topology is right. The benchmark is still someone else's marketing.

Sources
  1. AMD and Cerebras disaggregated inference partnership; "AMD Helios provides ultra-high throughput, processing prompts and large context windows" and "The Cerebras Wafer-Scale Engine accelerates the memory-bandwidth-intensive token generation"; the 5x T/s/W claim and its "Based on modelling by AMD Performance Labs and Cerebras in July 2026 ... comparing an AMD Helios rackscale solution with Cerebras WSE to a Cerebras WSE-only configuration" disclaimer. AMD Newsroom Jul 23, 2026
  2. Helios launch and the "up to 30% more inference tokens per dollar" claim, footnoted to "AMD Performance Labs estimates as of July 2026 ... using the Kimi K2 Thinking workload (32K input / 8K output) ... compared to an NVIDIA Vera Rubin NVL72 rack". AMD Investor Relations / AMD Newsroom Jul 23, 2026
  3. Cerebras side of the announcement, including deployment via Cerebras Cloud in H2 2026. Cerebras Jul 23, 2026