Skip to content

News, guides and free tools for the SaaS stack

LaunchLaunches

Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS for custom AI voices

The two text-to-speech models can design voices from a text prompt in 100+ languages and clone a voice from a 30-second sample, with SynthID watermarking on all audio.

Layered multicolored sound waveforms on a slate-blue background
Layered multicolored sound waveforms on a slate-blue background
Text size
Key takeaways
  • Gemini 3.8 Flash TTS targets creative voice work; Flash-Lite TTS targets high-volume, cost-efficient use.
  • Both are available now in the Gemini API and Google AI Studio; Gemini Enterprise support is coming soon.
  • Voice replication needs recorded consent and is unavailable in AI Studio in Illinois, Texas, the EEA, the UK, Switzerland and India.

Google has released two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, aimed at developers and companies that need custom, expressive synthetic voices. Announced on September 23, 2026, both are available today through the Gemini API and Google AI Studio.

Two models, two jobs

ModelBuilt for
Gemini 3.8 Flash TTSCreative direction and character design
Gemini 3.8 Flash-Lite TTSHigh-volume, cost-efficient applications

What the models can do

According to Google's announcement, the new models add:

  • Generative voice design: describe a voice in plain language and get a custom one, across 100+ languages and dialects.
  • A large voice library: 2,000+ production-ready voices, including regional varieties such as Mexican Spanish, Quebec French, Scots English and Brazilian Portuguese.
  • Voice replication: recreate a voice from a 30-second audio sample. This requires a recorded verbal consent check.
  • Line-by-line direction: control acting cues, pacing and emotional tone for each line.
  • Long-form audio: Google says quality holds across hours of continuous audio.
  • Two-speaker scenes: native support for directing conversations between two voices.
  • Vocal bursts and backchanneling: laughs, sighs, gasps and short interjections.

Google also lists voice remixing, which fine-tunes timbre, pitch, pace and accent, as coming soon.

Google's benchmark claims

Google says Gemini 3.8 Flash TTS ranks first on Hume AI's Voice Design Benchmark (71.4), and that the two models place first and second on Hume AI's Overall Quality Index. It also reports top positions on the Voice Arena leaderboard for several major languages, and improvements in long-form content and two-speaker control over Gemini 3.1 Flash TTS. These are Google's figures from its own announcement.

Where you can use it

ModelAvailable todayComing soon
Gemini 3.8 Flash TTSGemini API, Google AI Studio, Gemini NotebookGemini Enterprise
Gemini 3.8 Flash-Lite TTSGemini API, Google AI Studio, Google VidsGemini Enterprise

Developers can try the models in Google AI Studio. Google names Agora, LiveKit, Pipecat and Vercel as third-party integrations, and lists Figma, HeyGen, Linguana, Wondercraft, 99.co, Ollang, Spoken and Transforms.AI among the companies building with the models.

We couldn't confirm pricing from the announcement, so check the Gemini API documentation before budgeting for volume.

Safety and regional limits

All audio the models produce carries an imperceptible SynthID watermark so it can be detected as AI-generated, and Google says it uses C2PA credentials for provenance. Voice replication requires a recorded verbal consent.

Voice replication in AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland or India.

Why it matters

Speech is turning into a normal building block for SaaS products: onboarding walkthroughs, support agents, product videos and in-app narration. A cheaper Flash-Lite tier suits high-volume features, while the fuller Flash model targets branded or character voices. Teams that ship audio to users in the regions above should plan for the replication limits and for consent handling, and anyone already using voice tools built on other engines now has a new option to test on cost and quality.

Sources
#Google#Gemini#Text-to-speech#AI voice#Developers
Newsroom

The Appboxs newsroom covers launches, funding, acquisitions, pricing changes and AI across the SaaS and no-code world. Every story links to its primary sources. Have a tip, a correction or a story we should cover? Send it through our contact page.

Related stories

Launch

Anthropic launches Claude for Financial Services: 10 agents, 12 data connectors

Newsroom · 23 Sept 2026 · 3 min read
Funding

Ema raises $77M to replace SaaS busywork with teams of AI agents

Newsroom · 23 Sept 2026 · 2 min read
Analysis

AI agents are coming for the SaaS stack: five signals from September

Newsroom · 22 Sept 2026 · 2 min read
Analysis

MCP is turning no-code builders into the back end for AI assistants

Newsroom · 19 Sept 2026 · 2 min read