Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

Google's Gemini 3.5 Live Translate brings real-time voice translation to 70+ languages

The new audio model rolls out in Google AI Studio, Google Translate and Google Meet, with continuous speech-to-speech translation that stays a few seconds behind the speaker.

Dmytro Spodarets
Jun 11, 2026 · 2 min read

Google on June 9, 2026 introduced Gemini 3.5 Live Translate, its latest audio model for near real-time speech-to-speech translation across more than 70 languages, Google said in a blog post. The company says the model automatically detects the spoken language and generates smooth, natural-sounding speech that preserves the speaker's intonation, pacing and pitch.

Real-time speech translation is a long-standing target for AI assistants, where the difficulty is less raw accuracy than latency and natural delivery — keeping a translated voice flowing in step with a live conversation rather than in stilted, delayed chunks. Google says that, unlike turn-by-turn systems that wait for a speaker to finish, Gemini 3.5 Live Translate generates speech continuously and stays just a few seconds behind the speaker, delivering audio without awkward pauses.

Google said the model is rolling out starting June 9 across several of its products: for developers in public preview via the Gemini Live API and Google AI Studio; for enterprises in private preview in Google Meet for select business Google Workspace customers, with a broader rollout later in the year; and for general users through the Google Translate app on Android and iOS, credited to Google's Anuda Weerasinghe and Tony Lu. In Google Meet, Google said the model raises supported languages to 70-plus from a previous limit of five and enables more than 2,000 language combinations in a single meeting. On Android, a new "listening mode" lets users hear translations through the phone's earpiece.

The performance claims here are Google's own and are not independently benchmarked: descriptions of the translation as "fluid" and "natural" are the company's framing for a capability that real-world use across accents, noisy settings and less-common languages will test. Google said all audio generated by its models is watermarked with SynthID.

For users, the open questions are practical — how well the model holds up outside demos and how quickly the previews graduate to general availability. What is clear is that Google is putting its newest audio model to work first on translation, one of the most-used and most-demanding assistant tasks.


Dmytro Spodarets
Dmytro Spodarets
Founder & Editor-in-Chief

Founder and Chief Editor of Data Phoenix — a San Francisco Bay Area media and education platform focused on AI and Data.

More news