Hugging Face and Cerebras open-source a speech-to-speech pipeline running Gemma 4 at 1,851 tokens per second
Hugging Face and Cerebras released an open speech-to-speech AI pipeline running Google DeepMind's Gemma 4 31B at 1,851 tokens per second on Cerebras hardware to cut conversational lag.
Hugging Face and Cerebras on July 1, 2026 released an open, modular speech-to-speech pipeline that runs Google DeepMind’s Gemma 4 31B language model on Cerebras inference hardware. Cerebras says the setup reaches 1,851 tokens per second, about 35 times faster than a typical GPU endpoint. The stack is aimed at killing the multi-second pauses that make voice assistants feel broken.
The design bets that conversational voice AI fails on tail latency, not average speed. Rather than optimize the median response, the team targeted P95 tail latency, the slow 5 percent of responses that produce awkward stalls, and used Cerebras’s throughput to keep even those replies fast enough to feel like talk.
The pipeline is a chain of open components rather than a single model: Silero for voice-activity detection, Nvidia’s Parakeet-TDT for speech recognition, Gemma 4 31B on Cerebras for understanding and generating replies, and Alibaba’s Qwen3-TTS for the spoken output. Because each stage is swappable, developers can replace any model without rebuilding the whole system.
The stack is not a demo. It already powers more than 9,000 Reachy Mini robots in production, Hugging Face said, giving the release a real deployment rather than a benchmark chart. The work is credited to a group of Hugging Face contributors, among them Andres Marafioti, Amir Mahla and Leandro von Werra, alongside Cerebras.
The headline speed comes from Cerebras and describes its own hardware; the 1,851 tokens-per-second figure and the 35-times comparison are Cerebras’s own claim, not confirmed by Hugging Face’s post, and real conversational latency depends on the speech-recognition and text-to-speech stages as much as on the language model. What the release does offer is an open recipe others can run and measure themselves.
Founder and Chief Editor of Data Phoenix — a San Francisco Bay Area media and education platform focused on AI and Data.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
