Google Cloud Text-to-Speech

B
B tier on Text-to-Speech SoftwareScore 7.0 · #13 of 215
Android app
Not listed
Free plan
Yes
Paid plans from
$16/mo
Runs on
api, Web
cloud.google.com
The Google Cloud Text-to-Speech homepage

Summary

Google Cloud Text-to-Speech is an API for turning text or SSML into speech audio for applications. It offers more than 380 voices across more than 75 languages and variants, with requests sent through REST or gRPC and audio available in formats including MP3, LINEAR16, OGG_OPUS, MULAW, and ALAW. Gemini-TTS can produce single- or multispeaker speech, with natural-language prompts to guide style, accent, pace, tone, and emotional expression. Chirp 3 instant custom voice can make personalized voice models from as little as 10 seconds of audio and is available in more than 30 locales. Streaming supports low-latency conversations, while long-audio synthesis accepts up to 1 million bytes of input. The API also offers pitch, speaking-rate, and volume controls. Use cases listed include contact-center voicebots, device voice generation, and accessible electronic program guides. Billing must be enabled and counts input characters, including spaces, newlines, and SSML tags other than the mark tag. Pricing varies by voice model; Standard and WaveNet voices cost US$4 per 1 million characters, with 0 to 4 million characters free.

Who it is for

It suits developers and businesses adding generated speech to applications through an API. Its voice choices, speech controls, streaming, and custom-voice options serve a range of speech-generation needs.

What is good

  • More than 380 voices across 75+ languages and variants
  • REST and gRPC integration
  • Streaming supports low-latency conversations
  • Chirp 3 custom voices use as little as 10 seconds of audio

What to know first

  • Billing must be enabled to use the service
  • Billing counts spaces, newlines, and most SSML tags
  • Rates and free usage limits vary by voice model

Everything Xiaomi review

Google Cloud Text-to-Speech: the full review

Google Cloud Text-to-Speech combines a broad voice catalog with controls for speech style and delivery. Check the billing rules and the rate for the voice model you plan to use.

Google Cloud Text-to-Speech turns text or SSML into audio through an API, with both conventional voice options and newer prompted speech models. It suits developers building speech into apps, devices, or voicebots, especially when they need flexible voice controls and output formats.

Overview

The service spans more than 380 voices across more than 75 languages and variants, with REST and gRPC integration and streaming synthesis for real-time conversations. Its range is a practical advantage for products serving different markets or needing responsive spoken interaction; it is less compelling for someone who simply wants a ready-made listening app rather than an API.

There are several distinct billing approaches, from per-character voice rates to token-priced Gemini models. Billing must be enabled, and input character counts include spaces, newlines, and SSML tags other than the mark tag, so the effective cost depends on both the chosen model and how the text is formatted.

Key features

Voice choice and direction

Gemini-TTS can generate single- or multispeaker speech and accept natural-language direction for style, accent, pace, tone, and emotional expression. That gives developers more control over delivery than a voice choice alone, though the best fit depends on the model and its separate pricing. Chirp 3 instant custom voice can create a personalized voice model from as little as 10 seconds of audio and is available in more than 30 locales.

Generation controls and integration

The API supports streaming for low-latency conversations and long-audio synthesis for inputs up to 1 million bytes. Pitch can shift by up to 20 semitones, speaking rate can range from four times slower to four times faster than normal, and volume gain spans -96 dB to +16 dB. Output formats include MP3, Linear16, OGG Opus, MULAW, and ALAW, giving developers options for different delivery needs.

Security and support

Google Cloud maps security, privacy, and compliance controls to standards and uses independent certifications, attestations, and audit reports for verification. Documentation includes guides, API references, quotas and limits, release notes, and a support resource. These are useful foundations for production integration, but they do not remove the need to check quotas and billing against a particular workload.

Pricing

The service is usage-priced rather than a conventional seat-based subscription. Standard and WaveNet voices each show 0.00 USD per month, with 0 to 4 million characters free and a rate of US$4 per 1 million characters thereafter. Those are the lowest stated per-character rates and suit high-volume speech where those voice tiers meet the quality and control needs.

Chirp 3: HD voices show 0.00 USD per month, with 0 to 1 million characters free, then US$30 per 1 million characters. Studio voices also show 0.00 USD per month and 0 to 1 million characters free, then US$160 per 1 million characters. Instant custom voice shows 0.00 USD per month at US$60 per 1 million characters, with no free usage limit. These rates make the voice choice a significant cost decision: higher-end and personalized options cost more per character, and custom voice has no stated free allowance.

Gemini model billing is token-based and has no free usage limit stated. Gemini 2.5 Flash TTS and Gemini 2.5 Flash-Lite Preview TTS are priced at $0.50 per 1 million text tokens and $10.00 per 1 million audio tokens; Gemini 3.1 Flash TTS (Preview) at $1.00 and $20.00, respectively; Gemini 2.5 Pro TTS at $1.00 and $20.00; and Gemini 3.8 Flash TTS (Preview), through Dec 31, 2026, at $0.50 and $9.00. New customers can receive $300 in Google Cloud credits to try Text-to-Speech and other Google Cloud products. The credit is a trial opportunity, not a recurring allowance.

Platforms

Google Cloud Text-to-Speech is available through API and web access. Its REST and gRPC requests make it suited to software integration, while the documented voicebot, device speech, and accessible electronic program guide uses show the range of applications it can serve. Commercial use is supported.

Who it's for

Choose it if you are building an application or service that needs broad language and voice coverage, controllable speech delivery, streaming, or an API-based custom voice option. It is also a strong fit when audio format and integration choices matter. Look elsewhere if you want a consumer-focused reader or a simple fixed subscription: the model-specific usage rates and character-count billing require cost planning.

Pros and cons

  • Pros: More than 380 voices across more than 75 languages and variants give developers wide coverage.
  • Pros: Gemini-TTS prompting, extensive speech controls, streaming, and long-audio support suit varied and interactive applications.
  • Pros: Multiple export formats and REST or gRPC integration offer implementation flexibility.
  • Cons: Billing varies by voice or model, and input characters such as spaces and SSML tags count toward usage.
  • Cons: The API-centered offering is not aimed at people seeking a standalone text-to-speech listening app.

Alternatives

ElevenLabs is worth considering for a freemium API-and-web option with a free tier of 10,000 credits per month and three Studio projects; its Starter plan is 6.00 USD per month.

Typecast may suit users who want Android or iOS access alongside web and API support, with a free tier offering three projects, 3,000 lifetime credits, 1 GB storage, and 720p video exports.

Amazon Polly is another paid API-and-web service with a free plan; its Standard plan is 4.00 USD per month for Standard voices and speech or Speech Marks requests.

Fish Audio is an alternative with a free tier offering 8,000 monthly credits, up to seven minutes of generation, and up to 500 characters per generation, as well as self-hosted and desktop platform options.

Murf AI may be preferable for a web, API, or Windows workflow organized around projects: its free tier includes 10 projects and 10 minutes of voice generation, but no downloads.

Azure AI Speech offers a free tier of 500,000 characters per month and pay-as-you-go text-to-speech billed per character, making it a comparable API-and-web option for usage-priced deployment.

Speechify is an option for people who want Android, iOS, desktop, or web access; its free tier includes 10 voices and a maximum speed of 1.5x.

Inworld TTS is another freemium API-and-web alternative, with on-demand TTS-2 priced at $25 per 1M characters and TTS-2 Flash at $15 per 1M characters, plus free evaluation and prototyping.

For a broader comparison, browse AI Voice Generators, Text-to-Speech Software, or Game Voice Generators.

Verdict

Google Cloud Text-to-Speech is best for developers who need broad voice coverage, fine-grained delivery controls, and API-based speech generation. Its strongest reason to choose it is the combination of voice range, streaming, and flexible synthesis; its strongest reason to look elsewhere is the varied usage pricing, which calls for careful model and billing choices.

Google Cloud Text-to-Speech plans and pricing

All plans
Standard voices Free US$0.000004 per character (US$4 per 1 million characters) 0 to 4 million characters free cloud.google.com · 4 Oct 2026
WaveNet voices Free US$0.000004 per character (US$4 per 1 million characters) 0 to 4 million characters free cloud.google.com · 4 Oct 2026
Neural2 voices $16/mo Price after free usage limit is reached 1 million characters free usage cloud.google.com · 20 Sept 2026
Polyglot voices $16/mo Price after free usage limit is reached 1 million characters free usage cloud.google.com · 20 Sept 2026
Chirp 3: HD voices Free US$0.00003 per character (US$30 per 1 million characters) 0 to 1 million characters free cloud.google.com · 4 Oct 2026
Studio voices Free US$0.00016 per character (US$160 per 1 million characters) 0 to 1 million characters free cloud.google.com · 4 Oct 2026

Compared on text-to-speech software

Free plan
Nocloud.google.com
Voice cloning
Yescloud.google.com
Commercial use
Yescloud.google.com
API access
Yescloud.google.com

Facts

Export formats
MP3, LINEAR16, OGG_OPUS, MULAW, ALAWcloud.google.com · 20 Sept 2026
Platforms
Web, APIcloud.google.com · 20 Sept 2026
Purpose
The API converts text into natural-sounding speech and can accept text or SSML input to produce audio data.cloud.google.com · 4 Oct 2026
Voice selection
Google lists more than 380 voices across more than 75 languages and variants.cloud.google.com · 4 Oct 2026
Gemini-TTS
Gemini-TTS supports single- or multispeaker speech and lets users steer style, accent, pace, tone, and emotional expression with natural-language prompts.cloud.google.com · 4 Oct 2026
Custom voices
Chirp 3 instant custom voice can create personalized voice models with as little as 10 seconds of audio input and is available in more than 30 locales.cloud.google.com · 4 Oct 2026
Streaming
The product supports streaming audio synthesis for low-latency, real-time conversations.cloud.google.com · 4 Oct 2026
Long audio limit
Long audio synthesis supports up to 1 million bytes of input.cloud.google.com · 4 Oct 2026
Speech controls
The API supports pitch adjustment up to 20 semitones, speaking rates up to four times faster or slower than normal, and volume gain from -96 dB to +16 dB.cloud.google.com · 4 Oct 2026
Interfaces and formats
Applications can integrate through REST or gRPC requests, and output formats include MP3, Linear16, and OGG Opus.cloud.google.com · 4 Oct 2026
Use cases
Google describes use cases including contact-center voicebots, voice generation in devices, and accessible electronic program guides.cloud.google.com · 4 Oct 2026
Billing limit
Billing counts input characters including spaces, newlines, and SSML tags except the mark tag, and billing must be enabled to use the service.cloud.google.com · 4 Oct 2026
Security and compliance
Google Cloud says it maps security, privacy, and compliance controls to standards and undergoes independent verification through certifications, attestations, and audit reports.cloud.google.com · 4 Oct 2026
Trial credits
New customers can receive $300 in Google Cloud credits to try Text-to-Speech and other Google Cloud products.cloud.google.com · 4 Oct 2026
Support resources
The documentation page provides guides, API references, quotas and limits, release notes, and a getting-support resource.docs.cloud.google.com · 4 Oct 2026

Best Google Cloud Text-to-Speech alternatives

See all 20

Where it ranks on Everything Xiaomi

Is Google Cloud Text-to-Speech yours?

Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.

Sources