Best Speech-to-Text & Transcription APIs

Last verified 2026-06-21 · 14 picks · every field cited · no reviews yet · no paid placement

Opinionated picks for speech-to-text & transcription apis, with the trade-offs up front, judged from each API’s published, cited data. The reasoning is shown, so disagree where you know better. Compare all 14

Best overall: ElevenLabs Scribe (Speech to Text)

81 / 100 data score

ElevenLabs Scribe is a REST speech-to-text API supporting batch and real-time transcription across 90+ languages, with sub-150ms latency for streaming use cases. It covers speaker diarization, word and character timestamps, entity detection and redaction, multichannel processing, and keyterm prompting, making it suitable for podcasts, video captioning, meeting documentation, and AI agent integrations. Pricing starts at $0.22 per hour of audio with a free tier of 4.5 hours per month, self-serve signup, and an enterprise plan available. The service holds SOC 2 Type 2, HIPAA, GDPR, ISO 27001, and PCI DSS certifications, and ships SDKs for Python, Node.js, Swift, Kotlin, and Flutter.

From $0.22 hour of audio · free tier available.

How it stacks up

  • ElevenLabs Scribe (Speech to Text) starts lower at $0.22 hour of audio vs $1 for Azure AI Speech to Text.
  • ElevenLabs Scribe (Speech to Text) supports webhooks for event-driven integration; the other does not document them.
  • ElevenLabs Scribe (Speech to Text) supports webhooks for event-driven integration; the other does not document them.

Best for: Prototypes and side projects - free to start, no sales call; Regulated or enterprise workloads - compliance attestations and an enterprise plan; AI agents and automation - an agent-ready surface (MCP / llms.txt).

ElevenLabs Scribe (Speech to Text) profile → · ElevenLabs Scribe (Speech to Text) vs Azure AI Speech to Text

At a glance

14 picks ranked from published, cited fields
#APIBest forStarting priceScore
1ElevenLabs Scribe (Speech to Text)Best overall · Best free pick · Best for enterprise · Best for agents$0.22 hour of audio81
2Azure AI Speech to Texttransparent public pricing$1 hour of audio79
3Amazon Transcribetransparent public pricing$0.006 minute72
4Google Cloud Speech-to-Texttransparent public pricing$0.02 minute70
5IBM watsonx Speech to Texttransparent public pricing$0.02 minute73
6AssemblyAItransparent public pricing$0.0025 minute79
7Speechmaticstransparent public pricing$0.0022 minute67
8DeepgramBroadest surface - 59
9OpenAI Speech-to-Texttransparent public pricing$0.003 minute72
10Gladiatransparent public pricing$0.61 hour71
11Rev AItransparent public pricing$0.0017 minute58
12VoicegainCheapest to start$0.0015 minute60
13Sonioxtransparent public pricing$0.0017 minute68
14Groq Speech-to-Text (Whisper)transparent public pricing$0.04 hour of audio71

Quick pick by use case

If you only have thirty seconds, find your situation:

  • If you want the strongest all-round pick, pick ElevenLabs Scribe (Speech to Text) - our default pick: strongest across pricing, trust and breadth.
  • If you want to start free, pick ElevenLabs Scribe (Speech to Text) - free tier: Free plan includes 4 hours 30 minutes/month of Scribe v1/v2 transcription and 2 hours 30….
  • If you're buying for a regulated or large team, pick ElevenLabs Scribe (Speech to Text) - for regulated or large teams: SOC 2 Type II, HIPAA, enterprise plan.
  • If you want the lowest published entry price, pick Voicegain - from $0.0015 minute to start; compare on your real usage, not the entry price.
  • If you're wiring this into coding agents or AI workflows, pick ElevenLabs Scribe (Speech to Text) - easiest to wire up programmatically: MCP server + llms.txt.
  • If you want the broadest documented surface, pick Deepgram - 30 documented actions; breadth isn't quality, but it's the most to build on.

The picks in depth

  • #1 ElevenLabs Scribe (Speech to Text)

    81 / 100
    • Best overall
    • Best free pick
    • Best for enterprise
    • Best for agents

    ElevenLabs Scribe is a REST speech-to-text API supporting batch and real-time transcription across 90+ languages, with sub-150ms latency for streaming use cases. It covers speaker diarization, word and character timestamps, entity detection and redaction, multichannel processing, and keyterm prompting, making it suitable for podcasts, video captioning, meeting documentation, and AI agent integrations. Pricing starts at $0.22 per hour of audio with a free tier of 4.5 hours per month, self-serve signup, and an enterprise plan available. The service holds SOC 2 Type 2, HIPAA, GDPR, ISO 27001, and PCI DSS certifications, and ships SDKs for Python, Node.js, Swift, Kotlin, and Flutter.

    PricingHybrid · from $0.22 hour of audio · free tier
    TrustSOC 2 Type II · HIPAA · GDPR · ISO 27001 · PCI DSS
    Does
    • Real-time streaming
    • Speaker diarization
    • PII redaction
    Used byRevolut, Klarna, Washington Post, Deutsche Telekom
    Strengthstransparent public pricing · a free tier · self-serve signup
    ScoresAgent 65 · Pricing 100 · Setup 85 · Docs 55 · Procurement 100 · Trust 80

    ElevenLabs Scribe (Speech to Text) profile →

  • #2 Azure AI Speech to Text

    79 / 100

    Azure AI Speech to Text is Microsoft's cloud speech recognition service, offering real-time transcription, batch processing, speaker diarization, pronunciation assessment, and speech translation across more than 30 Azure regions. It starts at $1.00 per hour of audio with a free tier of 5 hours per month, scales via usage-based pricing, and supports self-serve signup with no sales call required. SDKs cover C#, Python, JavaScript, Java, Go, and Objective-C, and the service holds SOC 2 Type II, HIPAA, GDPR, ISO 27001, and PCI DSS certifications.

    PricingUsage · from $1 hour of audio · free tier
    TrustSOC 2 Type II · HIPAA · GDPR · ISO 27001 · PCI DSS
    Does
    • Real-time streaming
    • Speaker diarization
    • Speech translation
    Used byMicrosoft Teams, Microsoft Office 365, Microsoft Edge
    Strengthstransparent public pricing · a free tier · self-serve signup
    ScoresAgent 65 · Pricing 100 · Setup 85 · Docs 25 · Procurement 100 · Trust 100

    Azure AI Speech to Text profile →

  • #3 Amazon Transcribe

    72 / 100

    Amazon Transcribe is an automatic speech recognition service from AWS that converts audio to text via batch or real-time streaming, with support for speaker diarization, custom vocabularies, custom language models, and multi-language identification. It targets a broad range of applications including contact center analytics, clinical documentation through a dedicated medical variant, accessibility captioning, and toxic content detection in gaming. Pricing starts at $0.006 per minute on a pay-as-you-go basis, with a free tier of 60 minutes per month for the first 12 months. The service is HIPAA-eligible, SOC 2 Type 2 certified, ISO 27001 and PCI DSS compliant, available across 25 AWS regions including GovCloud, and provides SDKs for Python, JavaScript, Java, Go, C++, Ruby, and PHP.

    PricingUsage · from $0.006 minute · free tier
    TrustSOC 2 Type II · HIPAA · GDPR · ISO 27001 · PCI DSS
    Does
    • Real-time streaming
    • Speaker diarization
    • Medical transcription
    • PII redaction
    Strengthstransparent public pricing · a free tier · self-serve signup
    ScoresAgent 30 · Pricing 100 · Setup 85 · Docs 15 · Procurement 100 · Trust 100

    Amazon Transcribe profile →

  • #4 Google Cloud Speech-to-Text

    70 / 100

    Google Cloud Speech-to-Text is a REST API from Google Cloud that converts audio to text, supporting synchronous, batch, and streaming transcription across more than a dozen languages and regional endpoints. It covers call center transcription, live captioning with WebVTT and SRT output, speaker diarization, and multi-speaker meeting transcription. Pricing starts at $0.016 per minute with a free tier of 60 minutes per month, self-serve signup, and no sales call required. The service holds SOC 2 Type 2, ISO 27001, HIPAA, GDPR, and PCI DSS certifications, and ships official SDKs for Python, Node.js, Java, Go, C#, PHP, Ruby, and C++.

    PricingUsage · from $0.02 minute · free tier
    TrustSOC 2 Type II · HIPAA · GDPR · ISO 27001 · PCI DSS
    Does
    • Real-time streaming
    • Speaker diarization
    • Medical transcription
    Used byHubSpot, InteractiveTel, Embodied, iGenius
    Strengthstransparent public pricing · a free tier · self-serve signup
    ScoresAgent 20 · Pricing 100 · Setup 85 · Docs 15 · Procurement 100 · Trust 100

    Google Cloud Speech-to-Text profile →

  • #5 IBM watsonx Speech to Text

    73 / 100

    IBM watsonx Speech to Text is a REST API for fast, accurate transcription supporting batch, streaming, and WebSocket modes, aimed at customer self-service, call-center analytics, captioning, and accessibility applications. Pricing starts at $0.02 per minute with a 500-minute free tier and no sales call required, scaling to enterprise plans with unlimited concurrency. Deployments are available across seven global regions, SDKs cover Python, Node.js, Java, Swift, and Go, and the service holds SOC 2 Type II, HIPAA, GDPR, and ISO 27001 certifications.

    PricingUsage · from $0.02 minute · free tier
    TrustSOC 2 Type II · HIPAA · GDPR · ISO 27001
    Does
    • Real-time streaming
    • Speaker diarization
    Used byCitibank, Bradesco, Humana
    Strengthstransparent public pricing · a free tier · self-serve signup
    ScoresAgent 30 · Pricing 100 · Setup 85 · Docs 35 · Procurement 100 · Trust 90

    IBM watsonx Speech to Text profile →

  • #6 AssemblyAI

    79 / 100

    AssemblyAI is a voice AI platform providing speech-to-text transcription, speaker diarization, and audio intelligence features via REST API, aimed at developers building products on top of speech data. Pricing is usage-based at $0.0025 per minute with a $50 one-time free credit requiring no credit card, and enterprise plans are available. The service holds SOC 2 Type II, HIPAA, GDPR, ISO 27001, and PCI DSS certifications, with data processed in the US and EU. Customers include Zoom, Spotify, and Dovetail, and SDKs are actively maintained for Python and Node.js.

    PricingUsage · from $0.0025 minute · free tier
    TrustSOC 2 Type II · HIPAA · GDPR · ISO 27001 · PCI DSS
    Does
    • Real-time streaming
    • Speaker diarization
    • Speech translation
    • Medical transcription
    • PII redaction
    Used byZoom, Spotify, Veed, CallRail
    Strengthstransparent public pricing · self-serve signup · SOC 2 Type II
    Avoid ifYou want to try it free before paying
    ScoresAgent 70 · Pricing 85 · Setup 60 · Docs 75 · Procurement 85 · Trust 100

    AssemblyAI profile →

  • #7 Speechmatics

    67 / 100

    Speechmatics is a speech-to-text API supporting batch and real-time transcription across EU, US, and Australia regions, with capabilities including speaker diarization, language detection, translation, summarization, and audio event detection, making it suited for contact centers, legal, medical, and broadcast use cases. Pricing starts at $0.0022 per minute with a free tier of 3,000 minutes per month and self-serve signup, scaling to enterprise plans with dedicated regional endpoints. The API is REST-based with SDK support for Python, Node.js, .NET, and Rust, and holds SOC 2 Type 2, HIPAA, GDPR, and ISO 27001 certifications.

    PricingUsage · from $0.0022 minute · free tier
    TrustSOC 2 Type II · HIPAA · GDPR · ISO 27001
    Does
    • Real-time streaming
    • Speaker diarization
    • Speech translation
    • Medical transcription
    Used bywhat3words, 3Play Media, Veritone, Deloitte UK
    Strengthstransparent public pricing · a free tier · self-serve signup
    ScoresAgent 30 · Pricing 100 · Setup 85 · Docs 15 · Procurement 100 · Trust 70

    Speechmatics profile →

  • #8 Deepgram

    59 / 100
    • Broadest surface

    Deepgram provides real-time and batch APIs for speech-to-text, text-to-speech, and voice agents, plus audio intelligence features like summarization. Pricing is usage-based, published, and self-serve. It offers webhooks, four SDKs, and an official MCP server, with availability in North America and Europe. The platform carries SOC 2 Type 2, HIPAA, GDPR, and PCI DSS compliance with a published SLA.

    PricingUsage · free tier
    TrustSOC 2 Type II · HIPAA · GDPR · PCI DSS
    Does
    • Real-time streaming
    • Speaker diarization
    • PII redaction
    • Self-hosted option
    Strengthstransparent public pricing · self-serve signup · SOC 2 Type II
    Avoid ifYou want to try it free before paying
    ScoresAgent 40 · Pricing 60 · Setup 50 · Docs 45 · Procurement 75 · Trust 85

    Deepgram profile →

See all 14 Speech-to-Text & Transcription APIs compared →

How we rank

The headline score is the average of six 0-100 axes computed only from each API’s published, sourced fields: agent-friendliness, pricing transparency, setup speed, docs quality, procurement ease, and trust readiness. An unknown signal scores 0 for its axis - we credit what’s confirmed, never guess. The score is one input, not the verdict: we lead with each pick’s trade-off, and where a job has only one real option we say so rather than crown it. Full method on the methodology page.

Why trust apio

  • Every field cited. Each profile links the source for every claim - check us.
  • Public audit log. Every change to this data is recorded per field, with who changed it and why.
  • Published, deterministic methodology. The score is a formula over the same fields you can see - recompute it yourself.
  • Zero affiliate links, zero ads, zero paid placement. Money never moves rank.
  • No reviews yet - and we say so rather than synthesizing them.

Frequently asked questions

What is the best speech-to-text & transcription api?

ElevenLabs Scribe (Speech to Text) is our current top pick across pricing, trust, and developer-surface data (from $0.22 hour of audio). The right pick depends on your constraint: if you want the strongest all-round pick, ElevenLabs Scribe (Speech to Text); if you want to start free, ElevenLabs Scribe (Speech to Text); if you're buying for a regulated or large team, ElevenLabs Scribe (Speech to Text).

How are these Speech-to-Text & Transcription APIs ranked?

By a transparent data-readiness score computed from each API's published, sourced fields: pricing, free tier, self-serve access, compliance, webhooks/sandbox, and capability breadth. No reviews, no paid placement.

Which Speech-to-Text & Transcription APIs have a free tier?

ElevenLabs Scribe (Speech to Text), Azure AI Speech to Text, Amazon Transcribe, Google Cloud Speech-to-Text, IBM watsonx Speech to Text, Speechmatics, Gladia, Groq Speech-to-Text (Whisper).

See the full Speech-to-Text & Transcription APIs directory and each profile for the underlying data and citations, or compare the leaders: ElevenLabs Scribe (Speech to Text) vs Azure AI Speech to Text.