Google ships two Gemini speech models that clone a voice from 30 seconds
reported by 3 outlets· 6 articles · confidence: medium · first seen 2026-09-15 00:00 UTC
What this means for you
If you need synthetic speech, both models ship in the Gemini API and AI Studio now. The one cost figure available is 2.74 cents for 78 seconds of audio on Flash TTS; Flash-Lite is for high volume. Cloning needs a 30-second sample and passes consent verification, so confirm you hold the rights before designing a pipeline around it.
Google released two Gemini text-to-speech models, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, available through the Gemini API and Google AI Studio. Flash TTS designs voices from natural-language prompts across 100-plus languages, and MarkTechPost reports it leads Hume AI's Voice Design Benchmark at 71.4, a score with no evaluation date given. Both recreate a voice from a 30-second sample with consent verification, SynthID watermarking and C2PA credentials; the library holds 2,000-plus voices. Eight days earlier Google shipped Gemini 3.8 Live and Live Extended Thinking for voice agents, with no detail in the sources. Simon Willison measured 20 seconds and 2.74 cents to generate 78 seconds of audio.
Models in this story
Key facts
- ·Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are available through the Gemini API and Google AI Studio. source
- ·Gemini 3.8 Flash TTS creates voices from natural-language prompts across more than 100 languages and dialects. source
- ·Voice replication works from a 30-second audio sample and is paired with consent verification, SynthID watermarking and C2PA credentials. source
- ·Gemini 3.8 Flash TTS leads Hume AI's Voice Design Benchmark with a score of 71.4, reported without an evaluation date. source
- ·Simon Willison measured about 20 seconds and 2.74 cents to generate 1 minute 18 seconds of audio with Gemini 3.8 Flash TTS. source
- ·Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking were released on 15 September 2026 for voice agents; the available sources give no capability or pricing detail. source
What the sources say
- Google DeepMind Blog — Announcement page for the two Live voice-agent models; no body text available to us.
- Google DeepMind Blog — Vendor announcement covering custom voice creation, replication from short samples, language coverage and the safeguards attached.
- Simon Willison — Brief post on the Live audio release, with no readable body text.
- Simon Willison — Hands-on playground built on the API, with a measured generation time and cost for one clip.
- MarkTechPost — Trade coverage framing the Live release as infrastructure for production voice agents.
- MarkTechPost — Trade write-up noting availability, language coverage and the benchmark placement claimed for the design model.
Sources
The original reporting. Follow these — they did the work.
- Google DeepMind BlogIntroducing Gemini 3.8 Live and 3.8 Live Extended Thinking2026-09-15
- MarkTechPostGoogle Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents2026-09-15
- Simon WillisonGemini Live audio2026-09-15
- Google DeepMind BlogGemini 3.8 text-to-speech says hello2026-09-23
- Simon WillisonGemini 3.8 TTS Playground2026-09-23
- MarkTechPostGoogle Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design2026-09-23