Google Cloud Speech-to-Text
- Android app
- Not listed
- Free plan
- Yes
- Paid plans from
- $0.02/mo
- Runs on
- api, self-hosted, Web

Summary
Google Cloud Speech-to-Text converts audio into text and provides APIs for adding speech recognition to applications. It supports synchronous, asynchronous and streaming recognition for post-processing, periodic or real-time results, and lists support for 85+ languages and variants. Model adaptation lets users provide hints for specialized terms, rare words and phrases. Speaker diarization can identify who produced each utterance; the service also supports multichannel recognition and says it can handle noisy audio without additional noise cancellation. Specialized models are offered for voice control, phone calls and video transcription. API v2 includes data residency, audit logging and customer-managed encryption keys. Speech-to-Text On-Prem runs in customers’ private data centers and is offered through a sales contact. Developer access includes REST and RPC APIs, client libraries and command-line quickstarts. Billing depends on audio duration, channel count, recognition model, batch method and API version; each channel is billed separately. Some plans include 60 free minutes per month, while V2 standard recognition is listed at $0.016 / 1 minute for 0–500,000 minutes, with lower rates at higher monthly volumes.
Who it is for
This service suits developers adding speech recognition to applications and teams that need transcription in batch, periodic or real time. The on-premises option is for customers seeking to run it in private data centers.
What is good
- Supports synchronous, asynchronous and streaming recognition.
- Lists support for 85+ languages and variants.
- Can identify speakers and handle multichannel audio.
- API v2 includes data residency and audit logging.
- REST and RPC APIs, libraries and command-line quickstarts.
What to know first
- Each audio channel is billed separately.
- Pricing varies by duration, channels, model, batch method and API version.
- On-premises access is offered through a sales contact.
Everything Xiaomi review
Google Cloud Speech-to-Text: the full review
Google Cloud Speech-to-Text supports several recognition modes, speaker identification and specialized models, with billing determined by audio and processing choices. Check the applicable plan details: some include 60 free minutes per month, and charges vary by model and API version.
Overview
Google Cloud Speech-to-Text is an API service for turning audio into transcripts and adding speech recognition to applications. It is best suited to developers and organizations that need multiple ways to process audio, broad language coverage, or controls for sensitive workloads. Its flexibility is compelling, but the per-minute bill depends on model, channels, processing method, and API version.
Key features
Recognition can run synchronously, asynchronously, or as a stream, so teams can choose between immediate results, periodic jobs, and post-processing. Support for more than 85 languages and variants makes it useful for multilingual products, while model adaptation lets developers supply hints for specialized vocabulary and rare phrases. Models tuned for voice control, phone calls, and video transcription offer more targeted choices; that breadth is valuable, though it also means choosing the right model is part of managing cost and results.
Speaker diarization predicts which speaker produced each utterance, and timestamps plus VTT and SRT export support subtitle workflows. Multichannel recognition can preserve distinctions between channels, and the service says it can handle noisy audio without additional noise cancellation. These capabilities suit recordings with multiple speakers or challenging sound, but channel count matters to the bill because each channel is charged separately.
Developers can integrate through REST or RPC APIs, client libraries, and command-line quickstarts. API v2 provides data residency, audit logging, and customer-managed encryption keys for organizations with specific governance requirements. Speech-to-Text On-Prem runs in customers’ private data centers and is arranged through a sales contact, which may suit organizations that cannot use a cloud-hosted workflow.
Pricing
Pricing is usage-based rather than a simple flat subscription: charges vary with audio duration, channel count, model, batch method, and API version. The listed V1 Standard API plans include 60 free minutes per month. After that, Standard without data logging costs $0.024 / 1 minute, while Standard with data logging costs $0.016 / 1 minute. The lower rate makes the with-logging option cheaper for volume, while the without-logging plan costs more per minute; the choice should reflect the processing arrangement your application requires.
V2 Standard recognition models cost $0.016 / 1 minute for 0–500,000 minutes, with lower rates at higher monthly volumes. V2 Standard dynamic batch recognition costs $0.003 / 1 minute and is intended for lower-urgency processing, making it the economical choice when results do not need the standard recognition path’s urgency. The separate V2 Dynamic Batch listing shows $0.00 per month, but its billing terms specify $0.003 / 1 minute, so it is not a no-cost transcription plan.
Medical Dictation and Medical Conversation are each $0.08 per month after 60 free minutes per month. The listed V2 Standard recognition and V1 Standard entries also show monthly account prices of $0.02; because actual charges depend on usage, these figures should not be treated as a flat, unlimited subscription. The service has a free allowance, but sustained workloads should be budgeted by minute and channel rather than by the headline monthly amount.
Platforms
Speech-to-Text is offered through API, web, and self-hosted platforms. The API and developer tooling make it a fit for software teams embedding recognition; the on-premises option is for private data centers and requires contacting sales.
Who it's for
Choose it for application-level transcription, streaming recognition, multilingual workflows, speaker labeling, or subtitle generation, particularly when you need to tune recognition with vocabulary hints. It is also a strong candidate where API v2’s residency, audit, and encryption controls matter. It is less suitable for buyers seeking a straightforward per-seat desktop transcription app or predictable flat-rate pricing.
Pros and cons
- Pros: Synchronous, asynchronous, and streaming modes cover both live and batch workflows.
- Pros: More than 85 languages and variants, model adaptation, and specialized models support varied recognition needs.
- Pros: Speaker diarization, timestamps, and VTT/SRT exports help with multi-speaker and subtitle tasks.
- Pros: API v2 security controls and an on-premises option address organizational deployment requirements.
- Cons: Usage costs vary by model, API version, method, duration, and channel count, so bills require workload-aware planning.
- Cons: The cheaper dynamic-batch option is lower urgency, so it is not the fit for time-sensitive results.
Alternatives
Speechmatics is worth considering for a freemium alternative with API, web, desktop, Linux, and self-hosted platform coverage; its free offer includes $100 in credits, two concurrent real-time sessions, and 10 pre-recorded files per second.
AssemblyAI is another API-oriented option, with $50 in free audio credits, five new streaming connections per minute, and five concurrent pre-recorded transcriptions.
Notta is a better fit for someone seeking an end-user app across mobile, desktop, browser, and extension platforms: its free plan provides one seat and 120 transcription minutes per month.
Amazon Transcribe offers an API and web option with a free tier of 60 minutes per month for 12 months; consider it when that time-limited allowance suits your usage.
Happy Scribe has a free plan with unlimited meeting recordings, capped at 45 minutes per recording, plus a 10-minute AI trial.
Rev AI is an API alternative with English transcription plans including Whisper Fusion and Whisper Large.
MacWhisper suits Mac users who prefer a desktop app, with a free-forever plan for transcription and subtitle export and a one-time Pro plan priced at 64.00 EUR.
Transkriptor offers a Lite plan at 9.99 USD per month with 300 minutes, a straightforward option for individuals starting out.
For category comparisons, browse Speech-to-Text Software, Speech Recognition Software, Transcription Software, or Audio Transcription Software.
Verdict
Google Cloud Speech-to-Text is the right choice for developers who need flexible recognition modes, language breadth, and controls that can extend to private data-center deployment. Its chief advantage is the range of ways it can handle audio; its chief drawback is that pricing and processing choices demand careful workload planning. Look elsewhere if you want a simple end-user transcription app or an uncomplicated flat monthly price.
Google Cloud Speech-to-Text plans and pricing
All plansCompared on transcription software
- Free plan
- Nocloud.google.com
- Speaker identification
- Yescloud.google.com
- Timestamp support
- Yescloud.google.com
- Export formats
- VTT, SRTcloud.google.com
- API access
- Yescloud.google.com
Facts
- Purpose
- Speech-to-Text converts audio into text transcriptions and provides APIs for integrating speech recognition into applications.cloud.google.com · 3 Oct 2026
- Real-time and batch modes
- The service supports synchronous, asynchronous, and streaming speech recognition for post-processing, periodic, or real-time results.cloud.google.com · 3 Oct 2026
- Languages
- The product page states support for 85+ languages and variants.cloud.google.com · 3 Oct 2026
- Model adaptation
- Model adaptation lets users give hints to improve recognition of domain-specific terms, rare words, and phrases.cloud.google.com · 3 Oct 2026
- Speaker diarization
- The service can predict which speaker produced each utterance in a conversation.cloud.google.com · 3 Oct 2026
- Audio handling
- The service supports multichannel recognition and says it can handle noisy audio without additional noise cancellation.cloud.google.com · 3 Oct 2026
- Specialized models
- Google offers models tuned for uses including voice control, phone calls, and video transcription.cloud.google.com · 3 Oct 2026
- Security controls
- Speech-to-Text API v2 supports data residency, audit logging, and customer-managed encryption keys.cloud.google.com · 3 Oct 2026
- On-premises option
- Speech-to-Text On-Prem runs in customers’ private data centers and is offered through a sales contact.cloud.google.com · 3 Oct 2026
- Integration and developer access
- The documentation lists REST and RPC APIs, client libraries, and command-line quickstarts.docs.cloud.google.com · 3 Oct 2026
- Connected service example
- A Google Cloud tutorial describes using Speech-to-Text with Translation API to create localized video subtitles.cloud.google.com · 3 Oct 2026
- Billing limits
- Pricing depends on audio duration, number of channels, recognition model, batch method, and API version; each channel is billed separately.cloud.google.com · 3 Oct 2026
- Company
- Google Inc. was officially born in August 1998, and Google’s current headquarters is in Mountain View, California.about.google · 3 Oct 2026
Best Google Cloud Speech-to-Text alternatives
See all 20Where it ranks on Everything Xiaomi
Is Google Cloud Speech-to-Text yours?
Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.
Sources
- cloud.google.com/speech-to-text· checked 3 Oct 2026
- docs.cloud.google.com/speech-to-text/docs· checked 3 Oct 2026
- cloud.google.com/speech-to-text/pricing· checked 3 Oct 2026
- about.google/company-info/our-story/· checked 3 Oct 2026




