Modelwire
Subscribe

Google releases Gemini 3.8 text-to-speech with 2,000 voices and custom cloning

Illustration accompanying: Gemini 3.8 TTS Playground

Google shipped two new text-to-speech models built on Gemini 3.8, expanding its voice synthesis capabilities with access to over 2,000 pre-built voices and custom voice cloning from 30-second samples. The release signals intensifying competition in generative audio, where major labs are moving beyond static voice libraries toward personalized synthesis at scale. Simon Willison's playground tool lowers friction for developers to experiment with the API, potentially accelerating adoption patterns similar to early ChatGPT playground effects. Custom voice cloning at this accessibility level raises both product opportunity and policy questions around voice authentication and consent.

Modelwire context

Analyst take

The 30-second cloning threshold is the number worth scrutinizing. That's low enough to clone a voice from a single phone call or short podcast clip, which puts the consent and authentication questions in a different category than prior voice synthesis releases that required minutes of training audio.

Modelwire has no prior coverage to anchor this to directly, so context has to come from the broader competitive landscape. Google's move sits inside an accelerating pattern where frontier labs are treating audio as a first-class modality rather than an add-on. The mention of GPT-6 Astra in the story's entity list suggests this release is being positioned partly in response to OpenAI's voice roadmap. Willison's playground matters here because developer tooling has historically been where adoption races get decided early, and Google has sometimes lagged on that front even when its underlying models were competitive.

Watch whether third-party developers report meaningful voice fidelity at the 30-second threshold within the next 60 days. If cloning quality holds up on short samples in public testing, that sets a new floor for the category and pressures ElevenLabs and OpenAI to respond on sample-length requirements.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsGoogle · Gemini 3.8 · gemini-3.8-flash-tts · gemini-3.8-flash-lite-tts · Simon Willison · GPT-6 Astra

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. Simon Willison originally reported this story as Gemini 3.8 TTS Playground”. The full content lives on simonwillison.net. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Google releases Gemini 3.8 text-to-speech with 2,000 voices and custom cloning · Modelwire