Google shipped Gemini 3.8 Flash TTS and Flash-Lite TTS on 23 Sep 2026. If you already sell voiceovers — or quote Fiverr/Upwork “AI voice” gigs — you now see a free-feeling AI Studio playground, two API model IDs, and a wall of “expressive / 2,000+ voices / voice replication” claims. What you do not get from the launch post is a job → tier → stay-on-ElevenLabs answer before you quote the next sample.
This is a Flash vs Flash-Lite vs keep your current stack sheet. Not a voice-benchmark dump. Not a rewrite of Gemini Live voice agents — Live is real-time dialogue; this post is text-to-speech generation. Sibling voiceover and edit stacks stay on ElevenLabs for freelancers and Descript vs CapCut. Gemini subscription desks stay on Gemini AI Pro vs Ultra. Here the only decision is which TTS path earns the next 60–90 second client script.
Affiliate disclosure: some tool or platform links may later be affiliate or referral. Prices below are from Google’s live Gemini TTS blog and Gemini API pricing docs as of 28 Sep 2026 (SAST) — re-check before you quote a client. Not legal advice; not a guarantee of clients, savings, or income.
What shipped 23 Sep 2026 — Flash TTS + Flash-Lite TTS in plain English
On 23 Sep 2026, Google published Gemini 3.8 text-to-speech. Two generation models landed: gemini-3.8-flash-tts (creative / character / multi-speaker) and gemini-3.8-flash-lite-tts (high-volume / cost-efficient). Surfaces named on the post: Google AI Studio voice playground with a dual-speaker screenplay editor, the Gemini API, Gemini Notebook, Google Vids (Flash-Lite), and Flash Enterprise “coming soon” via API.
Google’s framing: generative voice design from natural-language prompts, an expanded library (blog claim: 2,000+ production-ready voices), voice replication from a short sample with consent verification, and line-by-line performance direction. Treat Hume / Voice Arena leaderboard claims on that post as Google’s marketing context — not Mr1Tech proof you put in a client proposal.
This release sits in the Gemini Audio family after 3.8 Live. Live = dialogue agents. This sheet = generation TTS. If you need the Live agent angle, that is a sibling topic — do not mash the two meters into one quote.
Flash vs Flash-Lite — pick by job, not by hype
Model docs put the split in a workload table, not a vibes chart.
Flash TTS (gemini-3.8-flash-tts): max voice fidelity, acting nuance, dialect coverage. Best for studio narration, complex multi-speaker dialogue, heavy vocal-burst acting, difficult pronunciation, regional dialects. Languages listed: 130.
Flash-Lite TTS (gemini-3.8-flash-lite-tts): high throughput, low latency, cost efficiency. Best for high-volume dubbing, read-aloud features, voice-agent cascades, everyday single-speaker generation, and voice replication at scale. Languages listed: 101.
Same API schema and prompting structure on both — so the chooser is workload, not “which SDK.” If the client sample needs two speakers arguing with stage directions, start on Flash in Studio. If the job is forty explainer takes with light tone control, start on Flash-Lite. Migration notes on the model docs matter if you still call gemini-3.1-flash-tts-preview: turn-level directions move into speech_metadata, and unary defaults return WAV with a header — do not double-wrap PCM.
SA English or Afrikaans client scripts are a packaging example for dialect prompting — not a claim that every accent is “solved.” Prefer one logged sample over a leaderboard screenshot in the proposal.
The price sheet freelancers actually need
Paid Standard rows on Gemini API pricing (TTS sections), through 31 Dec 2026:
- Flash TTS: $0.50 text in / $9.00 audio out per 1M tokens. - Flash-Lite TTS: $0.50 text in / $6.00 audio out per 1M tokens.
From 1 Jan 2027: Flash $1.00 / $18.00; Flash-Lite $1.00 / $12.00. Batch is about half those Standard audio/text rows on the same page. Free tier: free of charge where available; Google AI Studio is free of charge in available regions per pricing notes.
Audio tokens = 25 tokens per second of audio (pricing footnote). That formula is the only minute math you should show a client — for example, 60 seconds of audio ≈ 1,500 audio tokens before you multiply by the $/MTok audio-out row. Do not invent monthly “average freelancer TTS bills.”
Studio vs API desks
Google AI Studio is the try desk: voice design playground, dual-speaker screenplay editor, free in available regions. Use it for character samples and two-speaker demos before you open a production key.
Gemini API is the production desk: client-owned keys, spend caps, Batch for overnight dubbing queues. Skip API until you have a cap and a disclosure line in the SOW.
Google Vids is the path when the deliverable already lives as a Vids cut and Flash-Lite is the voice inside that workflow — not a reason to rebuild your whole stack around Vids.
Subscription Pro/Ultra context is complementary packaging on the Gemini Pro vs Ultra post — that sheet is not TTS token pricing.
Consent, watermarks, and regions that block voice replication
Google’s blog is explicit on trust rails. Voice replication requires consent verification — a verbal consent recording that matches the reference speaker. Every Gemini Audio clip carries SynthID watermarking; C2PA credentials are part of the transparency story. Disclose AI use to clients in the SOW; do not present a generated take as a human session unless the client asked for that packaging and local rules allow it.
Voice replication in AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland, and India (blog footnote). If you or the client sit in a blocked region, design/preset voices and licensed library paths still matter — replication is not a universal button. Marketplace gigs that assume “clone my voice in five minutes” need a geo check before you accept the brief.
Card and billing-country quirks for AI Studio / Gemini API are packaging only; do not invent ZAR or decline rates. This is contract hygiene, not legal advice.
When ElevenLabs (or Descript) still wins
Stay on ElevenLabs when the client’s brand voice is already cloned and approved there, marketplace buyers specify that stack, or your rights/consent workflow is already documented on that desk. Do not invent ElevenLabs plan prices here — re-check their live pricing if you quote that vendor. Cross-link: ElevenLabs for freelancers.
Stay on Descript (or a CapCut-led video desk) when the deliverable is edit-by-transcript video, not raw TTS stems. Cross-link: Descript vs CapCut. A cheaper Gemini take that forces a full re-edit in a foreign timeline is not a saving.
Switch or dual-desk only after one logged sample beats your current stack on rework time and the client accepts the disclosure and watermark reality. Fibre vs mobile for long batch renders is packaging — see fibre vs mobile data for AI work if uplink is the bottleneck.
Buy/skip matrix (sheet)
| Job | Flash (Studio/API) | Flash-Lite | Stay ElevenLabs | Vids-only | Skip API for now |
|---|---|---|---|---|---|
| Character / two-speaker client sample | Yes — Studio first | Weak | If brand voice already there | No | Until sample clears |
| Volume explainer / dubbing batch | Maybe for hero takes | Yes — first | If buyer requires it | If cut is Vids | If no spend cap |
| Client rights workflow already on EL | Only with change order | Same | Yes | — | — |
| Edit-by-transcript video deliverable | Stem only if needed | Stem only | Maybe | Maybe | Prefer Descript path |
| Replication from client sample | Check geo + consent | Same rails | If already approved there | — | If region blocked |
| No SOW disclosure line yet | — | — | — | — | Yes — write disclosure first |
Get the 1-page Flash / Flash-Lite / stay-ElevenLabs chooser
Optional printable of this free guide. Soft link until checkout. Not a “paid summary.” Not a guarantee of clients, savings, or income.
This week’s action — one 60–90s script, Flash and Flash-Lite
Take one non-secret 60–90 second client script. Generate once on Flash and once on Flash-Lite (or the Studio free path in an available region). Compare rework time and any bill meters. Note watermark / disclosure language you would put in the SOW. Fill one sheet row: Flash for character samples / Flash-Lite for volume / stay ElevenLabs when the client stack is already there / Vids-only if the cut lives in Google Vids / skip API until you have a spend cap.
If Studio quality wins but the client contract forbids undisclosed AI voice, the contract wins — not the playground. Sources: Google blog TTS post (23 Sep 2026), Gemini API pricing, Flash TTS model docs — re-check model IDs and Standard/Batch rows at paste time.
