Google Cloud Speech-to-Text

C
C tier on Transcription SoftwareScore 6.9 · #7 of 43
Android app
Not listed
Free plan
Yes
Paid plans from
$0.02/mo
Runs on
api, self-hosted, Web
cloud.google.com
The Google Cloud Speech-to-Text homepage

Summary

Google Cloud Speech-to-Text converts audio into text and provides APIs for adding speech recognition to applications. It supports synchronous, asynchronous and streaming recognition for post-processing, periodic or real-time results, and lists support for 85+ languages and variants. Model adaptation lets users provide hints for specialized terms, rare words and phrases. Speaker diarization can identify who produced each utterance; the service also supports multichannel recognition and says it can handle noisy audio without additional noise cancellation. Specialized models are offered for voice control, phone calls and video transcription. API v2 includes data residency, audit logging and customer-managed encryption keys. Speech-to-Text On-Prem runs in customers’ private data centers and is offered through a sales contact. Developer access includes REST and RPC APIs, client libraries and command-line quickstarts. Billing depends on audio duration, channel count, recognition model, batch method and API version; each channel is billed separately. Some plans include 60 free minutes per month, while V2 standard recognition is listed at $0.016 / 1 minute for 0–500,000 minutes, with lower rates at higher monthly volumes.

Who it is for

This service suits developers adding speech recognition to applications and teams that need transcription in batch, periodic or real time. The on-premises option is for customers seeking to run it in private data centers.

What is good

  • Supports synchronous, asynchronous and streaming recognition.
  • Lists support for 85+ languages and variants.
  • Can identify speakers and handle multichannel audio.
  • API v2 includes data residency and audit logging.
  • REST and RPC APIs, libraries and command-line quickstarts.

What to know first

  • Each audio channel is billed separately.
  • Pricing varies by duration, channels, model, batch method and API version.
  • On-premises access is offered through a sales contact.

Everything Xiaomi review

Google Cloud Speech-to-Text: the full review

Google Cloud Speech-to-Text supports several recognition modes, speaker identification and specialized models, with billing determined by audio and processing choices. Check the applicable plan details: some include 60 free minutes per month, and charges vary by model and API version.

Overview

Google Cloud Speech-to-Text is an API service for turning audio into transcripts and adding speech recognition to applications. It is best suited to developers and organizations that need multiple ways to process audio, broad language coverage, or controls for sensitive workloads. Its flexibility is compelling, but the per-minute bill depends on model, channels, processing method, and API version.

Key features

Recognition can run synchronously, asynchronously, or as a stream, so teams can choose between immediate results, periodic jobs, and post-processing. Support for more than 85 languages and variants makes it useful for multilingual products, while model adaptation lets developers supply hints for specialized vocabulary and rare phrases. Models tuned for voice control, phone calls, and video transcription offer more targeted choices; that breadth is valuable, though it also means choosing the right model is part of managing cost and results.

Speaker diarization predicts which speaker produced each utterance, and timestamps plus VTT and SRT export support subtitle workflows. Multichannel recognition can preserve distinctions between channels, and the service says it can handle noisy audio without additional noise cancellation. These capabilities suit recordings with multiple speakers or challenging sound, but channel count matters to the bill because each channel is charged separately.

Developers can integrate through REST or RPC APIs, client libraries, and command-line quickstarts. API v2 provides data residency, audit logging, and customer-managed encryption keys for organizations with specific governance requirements. Speech-to-Text On-Prem runs in customers’ private data centers and is arranged through a sales contact, which may suit organizations that cannot use a cloud-hosted workflow.

Pricing

Pricing is usage-based rather than a simple flat subscription: charges vary with audio duration, channel count, model, batch method, and API version. The listed V1 Standard API plans include 60 free minutes per month. After that, Standard without data logging costs $0.024 / 1 minute, while Standard with data logging costs $0.016 / 1 minute. The lower rate makes the with-logging option cheaper for volume, while the without-logging plan costs more per minute; the choice should reflect the processing arrangement your application requires.

V2 Standard recognition models cost $0.016 / 1 minute for 0–500,000 minutes, with lower rates at higher monthly volumes. V2 Standard dynamic batch recognition costs $0.003 / 1 minute and is intended for lower-urgency processing, making it the economical choice when results do not need the standard recognition path’s urgency. The separate V2 Dynamic Batch listing shows $0.00 per month, but its billing terms specify $0.003 / 1 minute, so it is not a no-cost transcription plan.

Medical Dictation and Medical Conversation are each $0.08 per month after 60 free minutes per month. The listed V2 Standard recognition and V1 Standard entries also show monthly account prices of $0.02; because actual charges depend on usage, these figures should not be treated as a flat, unlimited subscription. The service has a free allowance, but sustained workloads should be budgeted by minute and channel rather than by the headline monthly amount.

Platforms

Speech-to-Text is offered through API, web, and self-hosted platforms. The API and developer tooling make it a fit for software teams embedding recognition; the on-premises option is for private data centers and requires contacting sales.

Who it's for

Choose it for application-level transcription, streaming recognition, multilingual workflows, speaker labeling, or subtitle generation, particularly when you need to tune recognition with vocabulary hints. It is also a strong candidate where API v2’s residency, audit, and encryption controls matter. It is less suitable for buyers seeking a straightforward per-seat desktop transcription app or predictable flat-rate pricing.

Pros and cons

  • Pros: Synchronous, asynchronous, and streaming modes cover both live and batch workflows.
  • Pros: More than 85 languages and variants, model adaptation, and specialized models support varied recognition needs.
  • Pros: Speaker diarization, timestamps, and VTT/SRT exports help with multi-speaker and subtitle tasks.
  • Pros: API v2 security controls and an on-premises option address organizational deployment requirements.
  • Cons: Usage costs vary by model, API version, method, duration, and channel count, so bills require workload-aware planning.
  • Cons: The cheaper dynamic-batch option is lower urgency, so it is not the fit for time-sensitive results.

Alternatives

Speechmatics is worth considering for a freemium alternative with API, web, desktop, Linux, and self-hosted platform coverage; its free offer includes $100 in credits, two concurrent real-time sessions, and 10 pre-recorded files per second.

AssemblyAI is another API-oriented option, with $50 in free audio credits, five new streaming connections per minute, and five concurrent pre-recorded transcriptions.

Notta is a better fit for someone seeking an end-user app across mobile, desktop, browser, and extension platforms: its free plan provides one seat and 120 transcription minutes per month.

Amazon Transcribe offers an API and web option with a free tier of 60 minutes per month for 12 months; consider it when that time-limited allowance suits your usage.

Happy Scribe has a free plan with unlimited meeting recordings, capped at 45 minutes per recording, plus a 10-minute AI trial.

Rev AI is an API alternative with English transcription plans including Whisper Fusion and Whisper Large.

MacWhisper suits Mac users who prefer a desktop app, with a free-forever plan for transcription and subtitle export and a one-time Pro plan priced at 64.00 EUR.

Transkriptor offers a Lite plan at 9.99 USD per month with 300 minutes, a straightforward option for individuals starting out.

For category comparisons, browse Speech-to-Text Software, Speech Recognition Software, Transcription Software, or Audio Transcription Software.

Verdict

Google Cloud Speech-to-Text is the right choice for developers who need flexible recognition modes, language breadth, and controls that can extend to private data-center deployment. Its chief advantage is the range of ways it can handle audio; its chief drawback is that pricing and processing choices demand careful workload planning. Look elsewhere if you want a simple end-user transcription app or an uncomplicated flat monthly price.

Google Cloud Speech-to-Text plans and pricing

All plans
Speech-to-Text V2 Standard recognition $0.02/mo per 1 month / account Standard speech recognition cloud.google.com · 20 Sept 2026
Speech-to-Text V2 Dynamic Batch Recognition Free per 1 month / account Dynamic batch processing cloud.google.com · 20 Sept 2026
Speech-to-Text V1 Standard with data logging $0.02/mo 60 free minutes, then per 1 month / account 60 free minutes per month cloud.google.com · 20 Sept 2026
Speech-to-Text V1 Standard without data logging $0.02/mo 60 free minutes, then per 1 month / account 60 free minutes per month cloud.google.com · 20 Sept 2026
Medical Dictation $0.08/mo 60 free minutes, then per 1 month / account 60 free minutes per month cloud.google.com · 20 Sept 2026
Medical Conversation $0.08/mo 60 free minutes, then per 1 month / account 60 free minutes per month cloud.google.com · 20 Sept 2026

Compared on transcription software

Free plan
Nocloud.google.com
Speaker identification
Yescloud.google.com
Timestamp support
Yescloud.google.com
Export formats
VTT, SRTcloud.google.com
API access
Yescloud.google.com

Facts

Purpose
Speech-to-Text converts audio into text transcriptions and provides APIs for integrating speech recognition into applications.cloud.google.com · 3 Oct 2026
Real-time and batch modes
The service supports synchronous, asynchronous, and streaming speech recognition for post-processing, periodic, or real-time results.cloud.google.com · 3 Oct 2026
Languages
The product page states support for 85+ languages and variants.cloud.google.com · 3 Oct 2026
Model adaptation
Model adaptation lets users give hints to improve recognition of domain-specific terms, rare words, and phrases.cloud.google.com · 3 Oct 2026
Speaker diarization
The service can predict which speaker produced each utterance in a conversation.cloud.google.com · 3 Oct 2026
Audio handling
The service supports multichannel recognition and says it can handle noisy audio without additional noise cancellation.cloud.google.com · 3 Oct 2026
Specialized models
Google offers models tuned for uses including voice control, phone calls, and video transcription.cloud.google.com · 3 Oct 2026
Security controls
Speech-to-Text API v2 supports data residency, audit logging, and customer-managed encryption keys.cloud.google.com · 3 Oct 2026
On-premises option
Speech-to-Text On-Prem runs in customers’ private data centers and is offered through a sales contact.cloud.google.com · 3 Oct 2026
Integration and developer access
The documentation lists REST and RPC APIs, client libraries, and command-line quickstarts.docs.cloud.google.com · 3 Oct 2026
Connected service example
A Google Cloud tutorial describes using Speech-to-Text with Translation API to create localized video subtitles.cloud.google.com · 3 Oct 2026
Billing limits
Pricing depends on audio duration, number of channels, recognition model, batch method, and API version; each channel is billed separately.cloud.google.com · 3 Oct 2026
Company
Google Inc. was officially born in August 1998, and Google’s current headquarters is in Mountain View, California.about.google · 3 Oct 2026

Best Google Cloud Speech-to-Text alternatives

See all 20

Where it ranks on Everything Xiaomi

Is Google Cloud Speech-to-Text yours?

Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.

Sources