Skip to main content
Google’s highest-quality text-to-speech model, with inline vocal bursts and two-speaker dialogue. Released September 2026.

Model

Parameters

Vocal Bursts

Gemini 3.8 TTS places short non-speech sounds written in angle brackets at the exact point they appear in input:
Supported bursts include <laugh>, <sigh>, <cough>, <breath>, <gasp>, <short pause> and <long pause>. In two-speaker scripts, listener sounds go between pipes, like |mhm| or |yeah|. Bursts are performed, not spoken. The square-bracket tags of Gemini 3.1 Flash TTS ([whispering], [excited]) are not part of the 3.8 syntax. Use instructions for the overall delivery style instead.

Multi-Speaker Dialogue

Generate a conversation between up to 2 voices in one request. Two equivalent forms are accepted. Form 1 — dialogue array. Each line carries its own voice; lines are spoken in order:
speaker is optional — lines sharing a voice are treated as the same speaker automatically. If every line uses the same voice, the result is ordinary single-voice speech. Form 2 — script plus speakers map. Write the script in input with speaker names, and map each name to a voice:
Speaker names must match the names used in input exactly. Both forms are limited to 2 distinct voices — a third voice is rejected with a 400.

Voices

30 prebuilt voices. Use either the Gemini name or the OpenAI-style alias in voice and speakers[].voice:

Languages

Language is auto-detected from input (130+ languages). Set language only to force a specific locale:

Output Formats

Resources

List voices and languages from the Resources API:

Example

Response

Completed Response

Poll GET /v1/tasks/{task_id} until status is completed:
Volume discounts are available for almost all models. Reach out at contact@unifically.com to discuss custom pricing.