Google released new Gemini Live models today in the Gemini API and Google AI Studio, giving developers tools for building real-time, voice-first products. The release covers two speech-to-speech models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, and pairs them with Gemini 3.5 Transcribe, its speech-to-text model.
Gemini 3.8 Live is the company's native speech-to-speech model, and Google says it can carry on a conversation while performing tasks at the same time. For harder requests, Gemini 3.8 Live Extended Thinking is built to reason more deeply, and Google says it ranks first on Artificial Analysis' leaderboard. That ranking is the company's claim, not an independent result reported in the notice.
Gemini 3.5 Transcribe handles speech-to-text and supports more than 85 languages, according to Google. The model came out last month and posted an average Word Error Rate of 4.0% in streaming mode and 2.6% in non-streaming mode, the company says.
The announcement gives developers no pricing for any of the three models, and it does not say when the new Live models will reach general availability beyond their arrival in the API and AI Studio today. It also does not specify which languages Gemini 3.8 Live and 3.8 Live Extended Thinking support, a gap that matters for anyone building voice products outside English.
The launch follows Google's earlier push to put voice features in front of everyday users. In July, the company added a system-wide voice mode to Gemini for Mac that lets people hold the Fn key and dictate into any window, with filler words stripped and mid-sentence corrections caught. That update also included an opt-in screen-aware mode that lets the assistant see on-screen content to finish tasks, and it was rolling out globally in English, with more languages planned.
Today's release targets a different audience: developers wiring speech into their own applications rather than people dictating into the Gemini app.













