
Google Cloud Text-to-Speech
Google Cloud Text-to-Speech: Chirp voices and Gemini narration
Google Cloud Text-to-Speech adds multilingual narration and spoken responses to applications. It offers Chirp 3 HD voices, Gemini speech with natural-language delivery prompts, streaming audio and an approval-based custom-voice option.
Free plan availableAccount required
- Access
- Free and paid plans
- Usage pricing
- USD 4 / 1,000,000 charactersStandard / WaveNet
- Available on
- Web, API
Cloud speech for narration and interactive applications
Google Cloud Text-to-Speech provides several voice families through a managed speech API. Chirp 3 HD supplies multilingual stock voices for narration and spoken responses. Gemini TTS adds natural-language direction for tone, style, pace and emotional delivery, with single-speaker narration and model-specific dialogue support. Standard, WaveNet, Neural2 and Studio remain available with different capabilities and rates. Production use requires an authenticated, billing-enabled Google Cloud project.
Chirp 3 HD supports SSML and streaming speech. Its advanced pace, pause and pronunciation controls are in Preview, with language exceptions for some controls. Gemini uses delivery prompts rather than the same set of SSML settings. Cloud TTS can return formats such as MP3 and PCM, depending on the selected model and endpoint; Gemini through Vertex AI has a different raw-audio response. The current Cloud TTS guide lists Gemini 2.5 Flash and Pro alongside Preview Flash-Lite and 3.1 Flash. Gemini 3.8 speech is a separate Gemini Enterprise API Preview offering, even though it appears on the same pricing page.
Custom voice creation has a separate access requirement. Chirp 3 Instant Custom Voice needs sales allowlisting and recorded speaker consent in addition to a short reference clip. Character-priced families include model-specific monthly free usage, while Gemini charges separately for text input tokens and audio output tokens. SSML tags count towards character billing except mark tags. Google’s Cloud TTS logging guide says customer synthesis text and audio are not logged; custom voice recordings and other Gemini API routes have their own handling. Cloud terms also govern input permissions and use restrictions.
Best for
- Developers adding multilingual speech to applications
- Teams producing narration and spoken customer responses
- Products needing prompted delivery or streaming dialogue
Limitations
Billing is required and usage above free allowances is charged automatically. Language coverage, streaming formats and controls differ by model. Custom Voice requires sales approval and recorded consent. Chirp advanced controls and some Gemini models are Preview. Gemini 3.8 speech uses the separate Gemini Enterprise API. Cloud TTS data-logging statements do not describe every adjacent service.
Google Cloud speech-generation features
Chirp 3 HD stock voices
Multilingual voices and regional locales support narration, spoken responses and streaming applications.
Gemini delivery prompts
Direct tone, pace, accent and expression with natural-language instructions for the supported Gemini TTS models.
Single-speaker and dialogue audio
Gemini Flash and Pro support narration and multi-speaker speech; model and endpoint limits apply.
Model-specific pronunciation controls
Chirp supports SSML plus Preview pace, pause and custom pronunciation options, with locale-specific availability.
Approval-based custom voices
Create a Chirp voice from a short reference and prescribed recorded consent after sales allowlisting.
Character and token billing
Use character-priced voice families or separately metered Gemini input/output tokens, with cloud infrastructure billed separately.
Google Cloud Text-to-Speech pricing
Free plan available. Usage-based pricing. Visit Google Cloud Text-to-Speech for full plan details and current offers.
Example plans and usage rates; additional charges may apply.
Standard / WaveNet
USD 4 / 1,000,000 characters
Plan details
Per million characters beyond the listed 4-million monthly allowance; same pricing SKU, not two additive free quotas.
Neural2
USD 16 / 1,000,000 characters
Plan details
Per million characters beyond the listed 1-million monthly allowance.
Chirp 3 HD
USD 30 / 1,000,000 characters
Plan details
Per million characters beyond the listed 1-million monthly allowance.
Technical specifications
Voice generation
| Feature | Google Cloud Text-to-Speech |
|---|---|
| Speech languages | Multilingual voices and regional locales across Chirp, Gemini and legacy families. Language coverage and Preview availability depend on the selected model; custom-voice locale support is separate.Model-specific language and regional availability |
| Voice and pronunciation controls | Chirp 3 HD supports SSML and Preview pace, pause and pronunciation controls. Pace spans 0.25–2×; pause and pronunciation controls have locale exceptions.Advanced controls Preview; availability varies by locale |
| Voice cloning | Chirp 3 Instant Custom Voice requires sales allowlisting, a prescribed recorded-consent statement and a reference clip, each up to 10 seconds. Supported languages and regions are limited.Sales approval required; recorded speaker consent |
| Audio exports and streaming | Cloud TTS Gemini unary requests support LINEAR16, PCM, MP3, Ogg Opus and A-law/μ-law. Streaming supports PCM, Ogg Opus and telephony encodings. Vertex AI returns raw 16-bit 24 kHz PCM.Encoding and streaming support depend on API and model |
| Commercial-use terms | Google Cloud AI terms treat generative output as Customer Data and do not assert ownership of newly created output IP. Service restrictions and rights to supplied text or voice recordings still apply.Cloud agreement and applicable use restrictions |
| API access | SupportedBilling-enabled Cloud project and authentication required |
Google Cloud Text-to-Speech alternatives
Other services offer AWS integration, managed streaming voices or creator-focused speech tools.
Amazon Polly
Four AWS speech engines with SSML, lexicons and per-character billing.
Compare with Google Cloud Text-to-SpeechCartesia
Sonic streaming speech and custom voices with shared speech credits.
Compare with Google Cloud Text-to-SpeechElevenLabs
Hosted speech and cloning with a browser studio and other audio production tools.
Compare with Google Cloud Text-to-SpeechGoogle Cloud Text-to-Speech FAQs
Is Google Cloud Text-to-Speech free?
Selected character-priced families include monthly free usage, but billing must be enabled and excess usage is charged. The listed allowance is 4 million characters for Standard/WaveNet and 1 million for Neural2, Chirp 3 HD and Studio. Gemini and Instant Custom Voice have no listed free allowance.
How is Gemini text-to-speech billed?
Text input and audio output are metered separately in tokens. Gemini 2.5 Flash costs USD 0.50 per million input tokens plus 10 per million output tokens; Pro costs 1 plus 20. Audio corresponds to 25 tokens per second. Other models have different rates and availability.
Can Google Cloud clone my voice?
Chirp 3 Instant Custom Voice can create a personal voice, but access requires sales allowlisting. It needs a prescribed recorded-consent statement and a reference recording, each up to 10 seconds. Supported languages and regions are limited.
Is Gemini 3.8 speech available through the Cloud TTS API?
The current documentation places Gemini 3.8 Flash and Flash-Lite speech in Preview through the Gemini Enterprise API only. Cloud TTS supports its documented Gemini 2.5 models and 3.1 Flash Preview; their controls, prices and endpoint support differ.
Is this your product? Claim this page to update your listing.