Gemini can copy a voice from 30 seconds of audio. Not in Illinois
2026-09-23 · 4 min read
Google DeepMind shipped two new text-to-speech models on Wednesday morning. @GoogleDeepMind introduced Gemini 3.8 Flash TTS, for designing voices with their own accents and character, and Gemini 3.8 Flash-Lite TTS, built for volume. About 15 minutes later Google's Logan Kilpatrick, @OfficialLoganK, listed the highlights: a new voice design tool, "2,000+ production ready voices, voice replication, support for 100 languages," and the top spot on Hume AI's voice benchmarks.
Voice replication is the feature most people will try first. It builds a copy of a real person's voice from a short recording. It's also the one feature a business here in St. Charles can't get through Google's AI Studio. A single line in Google's announcement says replication there "is not available in Illinois, Texas, EEA, UK, Switzerland, and India."
What Google shipped
Voice design starts from a written description. The model builds the voice and saves it to your project so you can reuse it. Google's docs cap that at 200 stored voices per project, each kept for a year.
The bigger change is direction. In a follow-up post, @GoogleDeepMind says you can adjust delivery line by line, down to pacing, emotion, and cues like a laugh or a pause. The blog adds two-speaker scenes and long recordings that stay consistent for hours.
Replication has a real gate on it. According to Google's voice replication docs, you supply a 10 to 30 second clip of the speaker, then a second recording of the same person reading a fixed consent statement word for word, saying they own the voice and agree to Google building a synthetic model of it. The system checks that the two voices match before it creates anything. The docs even suggest recording both on the same microphone in the same room so the check passes.
Every clip also gets a watermark. @GoogleDeepMind says "All generated audio is watermarked with SynthID so it can be reliably identified as AI-generated."
The price, and the fine print
Kilpatrick's second post says the new models cost less than the older 3.1 Flash TTS. Google's pricing page backs that up, with a catch. Flash TTS audio is $9 per million output tokens through December 31, then $18 starting January 1, 2027. The 3.1 preview model charged $20. Google counts 25 tokens per second of audio, so an hour of speech costs about 81 cents this year and about $1.62 next year. Flash-Lite is $6, rising to $12.
So the launch price is a half-off promotion, and the January price comes in about 10% under the old model. For a small business the difference barely registers, since an hour of narration costs less than a coffee at either price.
One caution on quality claims: the benchmark numbers come from Google's own announcement, and I haven't seen independent listening tests yet. Listen to samples in your own use case before you trust a leaderboard.
Why Illinois is on the list
Google doesn't say why. My guess is biometric privacy law. Illinois' Biometric Information Privacy Act names voiceprints specifically and lets individuals sue, and Texas has its own biometric law that covers voiceprints too. A tool that matches one recording of your voice against another is working with exactly that kind of data. Leaving those places out of AI Studio looks like the cautious legal call.
What it means for a business
Most of the useful parts still work in Illinois. The voice library, voice design, line-by-line direction and the language range cover nearly everything a small business would do with a synthetic voice:
- A phone agent or after-hours line that sounds calm instead of robotic
- On-hold messages and appointment reminders in English and Spanish, with regional varieties like Mexican Spanish in the library
- Narration for training videos and how-to clips, especially since Flash-Lite TTS is also coming to Google Vids
LiveKit and Pipecat, two common toolkits for building voice agents, are among the launch partners, so expect phone-bot vendors to adopt these voices quickly.
Cloning is the part to handle carefully. If a vendor offers to copy your voice or an employee's, ask where the sample is stored, what consent they collect, and how you delete it later. Illinois employers have already been sued under BIPA over fingerprint time clocks, and a voice is the same kind of data. The watermark also has a limit: it helps identify Google's audio when someone checks, but it won't stop a scam call made with a different tool.
A better voice also doesn't mean better answers. Last week a budget Gemini voice model scored 30% on customer-service tasks. What the bot knows and when it hands off to a person still matter more than how it sounds.
If you're weighing a voice agent for missed calls, our free process audit starts with how calls reach you today and where they get dropped.