Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

AMD and Cerebras partner on disaggregated AI inference, claiming 5x efficiency gain

AMD and Cerebras unveiled a disaggregated AI inference partnership claiming up to five times higher tokens per second per watt, launching first on Cerebras Cloud in late 2026.

D
Jul 23, 2026 · 1 min read

AMD and Cerebras Systems announced a partnership for disaggregated AI inference on July 23, 2026 at AMD’s Advancing AI event. The companies claim the combined setup delivers up to five times more tokens per second per watt than a Cerebras-only configuration.

The design splits the work. AMD’s Helios rack-scale infrastructure — EPYC server processors paired with Instinct MI400-series accelerators — handles high-throughput prompt processing, while Cerebras’ Wafer-Scale Engine handles the memory-bandwidth-heavy job of generating tokens at low latency.

Splitting inference into a prompt-processing stage and a token-generation stage lets each run on the silicon best suited to it, the companies said. The combined offering is expected first through Cerebras Cloud in the second half of 2026.

The 5x efficiency figure is a vendor claim measured against Cerebras’ own hardware and has not been independently verified; real-world gains depend on model, batch size and workload.

Separately at the same event, AMD detailed its 2027 server-CPU roadmap: EPYC “Venice-X” chips built on the Zen 6 architecture, plus “Verano” parts, with two-socket Venice-X platforms reaching up to 192 cores and 384 threads at more than 5GHz, aimed at agentic-AI and high-performance-computing workloads.

More news