Cerebras unveils CS-4 wafer-scale inference system, claiming up to 30x GPU speed
Cerebras unveiled the CS-4, a wafer-scale inference system combining three WSE-3T processors, and said it runs large-model inference up to 30 times faster than GPU-based systems.
Cerebras unveiled the CS-4, a wafer-scale inference system that packs three next-generation processors into a single server, at its Supernova event on Aug. 19. The company said it runs large-model inference up to 30 times faster than GPU-based systems.
Each CS-4 combines three Wafer Scale Engine 3 Turbo (WSE-3T) chips — each roughly the size of a dinner plate, with about four trillion transistors and 900,000 AI-optimized cores — for a combined 750 petaflops of sparse FP16 compute, according to Cerebras. The design targets the largest frontier models: the company says the CS-4 supports models above 50 trillion parameters and delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters.
The bet is that inference, not training, is where wafer-scale hardware pays off, as AI providers race to serve ever-larger models at usable speeds. Built on Cerebras’ Nexus server architecture, the CS-4 is up to twice as fast as the prior CS-3, uses 50% fewer components and runs a wafer-to-wafer interconnect at 2-microsecond latency, the company says.
Those numbers are Cerebras’ own and have not been independently benchmarked. The 30x comparison depends on which GPU systems, models and configurations are used, and vendor speed claims routinely narrow under third-party testing.
First CS-4 shipments are planned for this quarter, which will show whether the throughput holds up outside Cerebras’ own demos.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
