Gemini TTS (Google)
by Google · AI Voice
Current version: Gemini 3.8 Flash TTS (gemini-3.8-flash-tts; model page 'Latest update: September 2026') [official source]
Last tested: never · Facts last updated: Oct 8, 2026 · Official pages checked: Oct 8, 2026
What it is
Google's text-to-speech models (Gemini 3.8 Flash TTS) with voice design and consent-based voice replication.
Added 2026-10-08 as the round-1 voice replacement for Hume (Hume's TTS API shuts down 2026-11-13). Official prices are listed 'through December 31, 2026'; from 2027-01-01 they are $1.00 per 1M text input tokens and $18.00 per 1M audio output tokens (Flash).
Scores by job
A tool can be #1 for one job and mediocre for another, so we score each job separately (0–100).
- NOT YET TESTED
- NOT YET TESTED
- NOT YET TESTED
- NOT YET TESTED
- NOT YET TESTED
Test results
Objective metrics recorded by our test harness (same inputs for every tool). Raw outputs are archived as evidence.
Subjective assessment by our evaluators, blind where possible. This is opinion, labelled as such — not a measured fact.
Evidence: Gemini TTS (Google)'s raw outputs
Same input for every tool. We publish the first output (no cherry-picking), with the date, plan and model version used, plus a SHA-256 hash of the file.
When the test round runs, each task appears here as a row with one column per tool.
Best at
EditorialNOT YET TESTED
Pros & cons
EditorialNOT YET TESTED
Who should use it
NOT YET TESTED
Who should not
NOT YET TESTED
Pricing
| Plan | Price | Includes | Checked |
|---|---|---|---|
| Free Tier (Gemini API / AI Studio) | Free | rate-limited; content may be used to improve Google products | Oct 8, 2026 |
Official API prices (pay-as-you-go, from the vendor's developer docs)
| Model / option | Price (USD) | Checked |
|---|---|---|
| gemini-3.8-flash-tts audio output (standard) · audio tokens; 25 tokens per second; official: 'Equivalent to $0.00225 per 10s audio' (through 2026-12-31) | $9 / 1M output tokens | Oct 8, 2026 |
| gemini-3.8-flash-tts audio output (standard) · derived from the official '$0.00225 per 10s audio' | $0.000225 / sec | Oct 8, 2026 |
| gemini-3.8-flash-tts text input (standard) · through 2026-12-31 | $0.5 / 1M input tokens | Oct 8, 2026 |
| gemini-3.8-flash-tts audio output (batch/flex) | $4.5 / 1M output tokens | Oct 8, 2026 |
| gemini-3.8-flash-lite-tts audio output (standard) · official: 'Equivalent to $0.0015 per 10s audio' | $6 / 1M output tokens | Oct 8, 2026 |
Privacy & data handling (official-page facts, not a score)
Quoted from Gemini TTS (Google)'s official policy, security and pricing pages. Our privacy score (methodology v0.2) is published only after the full desk check and in-account verification. Until then: NOT YET TESTED. Items not listed here are not stated on the pages we checked.
- Training on your data:Paid Tier: not used to improve products; Free Tier / AI Studio: used, human review possiblepaid
“Google doesn't use your prompts (including associated system instructions, cached content, and files such as images, videos, or documents) or responses to improve our products”
source (official page, archived text) · checked Oct 8, 2026 · not yet verified in a test
- Training on your data:Unpaid services: content used to improve Google productsfree
“Google uses the content you submit to the Services and any generated responses to provide, improve, and develop Google products”
source (official page, archived text) · checked Oct 8, 2026 · not yet verified in a test
- Retention & deletion:Paid: prompts/responses logged for a limited period, for abuse detection onlypaid
“Google logs prompts and responses for a limited period of time, solely for detecting and preventing violations of the Prohibited Use Policy”
source (official page, archived text) · checked Oct 8, 2026 · not yet verified in a test
- GDPR / DPA / residency:Paid: processed under Google's Data Processing Addendumpaid
“will process your prompts and responses in accordance with the Data Processing Addendum for Products Where Google is a Data Processor”
source (official page, archived text) · checked Oct 8, 2026 · not yet verified in a test
- Participant notice / clone consent:voice replication requires consent audio
“Replicate a speaker's voice from reference and consent audio”
source (official page, archived text) · checked Oct 8, 2026 · not yet verified in a test
Features & limits (vendor-stated, not verified by our tests)
- Single- and multi-speaker speech
- Voice design from a natural-language description
- Voice replication from reference and consent audio
- Two models: 3.8 Flash TTS (fidelity) and 3.8 Flash-Lite TTS (low latency / cost)
- API / AI Studio only: no consumer voiceover editor
- Paid prices double on 2027-01-01 per the official pricing page
- Custom voices: 200 per project; stateful voices expire 1 year after last use