AMD entered a technical partnership with Cerebras to build a disaggregated AI inference platform combining Helios rackscale systems with Cerebras Wafer-Scale Engine.
Design targets ultra-low-latency token generation with high-throughput prompt processing; companies project up to 5x higher tokens per second per watt.
Cerebras plans to deploy Helios in its data centers; initial availability expected via Cerebras Cloud in 2H 2026.