Menu
≈ why?
See the rankings
← All platforms

Speechmatics

Speech-to-text Free tier

Enterprise speech-to-text with very broad language coverage and real on-prem options, for teams who self-host.

Best for wide language coverage, or running speech-to-text on your own hardware
Watch for one building block to wire up, not a finished agent
Free to try 3,000 min/mo STT + 1M TTS chars · no card

Paid link, we may earn a commission. How this works.

Our scores editorial preview
4.9 Fair overall / 10
Voice quality 3
Voice range 4
Ease of use 5
Value 8
All-in /min $0.01–0.01
headline /min $0.01
✓ HIPAA✓ SOC 2 Type II✓ GDPR

Scored on the same voice-agent rubric as the full platforms, so a building block like this scores low on the axes it does not address. Read its value score against its job.

See how it stacks up · Full rankings →

The languages-and-deployment specialist. Speechmatics turns speech into text in 56+ languages and will run inside your own data centre, not just its cloud. It is one building block though, not a whole phone agent. No voice, no language model, no phone line.

What you'll pay

About $0.01 to 0.01 for a minute of conversation, once the phone line and the AI are added in.

That's roughly $0.36–0.64 an hour. Plans: $0/mo (Free).

Pricing

$ 0.01–0.01/min The total you actually pay for one minute of conversation once every piece is added up: the platform, the AI, the voice and the phone line. ≈ €0.01–0.01≈ £0.00–0.01≈ ₹0.57–1.02≈ R$0.03–0.05≈ A$0.01–0.02 headline $0.01 /min
Show the cost breakdown
What the platform charges to run the agent, before the phone line and the AI usage are added on.
The step that turns what the caller says out loud into text the AI can read. $0.01 /min
The AI 'brain' that reads what the caller said and works out what to say back.
The step that turns the AI's written reply back into a spoken voice.
The phone line itself: the service that connects the call to a real phone number. Usually billed on top of the platform.
The total you actually pay for one minute of conversation once every piece is added up: the platform, the AI, the voice and the phone line. $0.01–0.01 /min

Speechmatics prices per HOUR of audio, not per minute, and the figures stored here are LIST rates, pixel-verified 2026-07-23 by re-reading the JS-rendered table with the 'Model Training: enable for 33% discount' toggle switched OFF: real-time STT $0.36/hr Standard and $0.64/hr Enhanced, batch $0.36/hr Standard and $0.60/hr Enhanced, so the per-minute band here is $0.006 (Standard) to about $0.0107 (Enhanced). Mind the page's default state: the table renders with that opt-in Model Training data-sharing discount ON, displaying $0.24/$0.43 real-time and $0.24/$0.40 batch, which is what this profile stored until this correction. Melia 1 (launched 2026-06-17), a batch-only production preview, lists at $0.192/hr; the advertised 'from $0.129/hr' (also the Pro plan's headline price) is the discounted figure. Text-to-speech is $0.011 per 1,000 characters (English at launch; the same figure with the toggle on or off), and the newly published bolt-ons (translation $0.65/hr, summaries $0.12/hr, chapters $0.40/hr, sentiment $0.12/hr, topics $0.20/hr) also read the same in both toggle states. The free tier is 3,000 minutes a month (50 hours), split 1,200 real-time plus 1,800 batch, plus about 1M free TTS characters, no card needed. A separate 20% volume discount applies above 500 hours a month per service. This is primarily speech-to-text: there is no language model and no telephony, so those components are 0 here. To run a full phone agent you add an LLM and a phone line separately, each a cost on top.

Plans & what you get

Every plan in one place: the monthly fee, what each one includes, and the features it unlocks. Anything beyond a plan's allowance, or on a pay-as-you-go tier, is billed at the per-minute rate above. A blank in the features means the vendor's plan page does not state it for that plan, not that it is unavailable.

FreeProEnterprise
Price FreeCustom
Included 3,000 minutes Pay per use
Plan notes 3,000 free minutes (50 hours) per month, split 1,200 real-time + 1,800 batch, no card required, plus ~1M free TTS characters; 2 concurrent real-time sessionsPay-as-you-go on usage. LIST rates: real-time STT $0.36/hr Standard, $0.64/hr Enhanced; batch $0.36/hr Standard, $0.60/hr Enhanced; TTS $0.011 per 1,000 characters (a rate the discount toggle does not touch). The rate table renders with the opt-in Model Training data-sharing discount (about 33%) ON by default, displaying $0.24/$0.43 real-time and $0.24/$0.40 batch instead. Melia 1 multilingual model in batch-only production preview: $0.192/hr list, advertised 'from $0.129/hr' with the discount applied. Capped at 6,000 hours/month, 50 concurrent real-time sessions.Custom pricing, no rate limits, on-prem/container deployment, volume discounts from 24,000 hours/year
What each plan unlocks
API access Yes Yes
Concurrent calls 2 real-time sessions 50 real-time sessions
Priority support Custom deployment + volume pricing
  • Free Free
    3,000 minutes

    3,000 free minutes (50 hours) per month, split 1,200 real-time + 1,800 batch, no card required, plus ~1M free TTS characters; 2 concurrent real-time sessions

    API access
    Yes
    Concurrent calls
    2 real-time sessions
    Priority support
  • Pro
    Pay per use

    Pay-as-you-go on usage. LIST rates: real-time STT $0.36/hr Standard, $0.64/hr Enhanced; batch $0.36/hr Standard, $0.60/hr Enhanced; TTS $0.011 per 1,000 characters (a rate the discount toggle does not touch). The rate table renders with the opt-in Model Training data-sharing discount (about 33%) ON by default, displaying $0.24/$0.43 real-time and $0.24/$0.40 batch instead. Melia 1 multilingual model in batch-only production preview: $0.192/hr list, advertised 'from $0.129/hr' with the discount applied. Capped at 6,000 hours/month, 50 concurrent real-time sessions.

    API access
    Yes
    Concurrent calls
    50 real-time sessions
    Priority support
  • Enterprise Custom

    Custom pricing, no rate limits, on-prem/container deployment, volume discounts from 24,000 hours/year

    API access
    Concurrent calls
    Priority support
    Custom deployment + volume pricing

Each plan bundles a set amount of talk time a month.

Prices in USD as set by the vendor · last checked 2026-07-30 · vendor pricing →

At a glance

· Plugging in your own phone-number supplier instead of using the platform's numbers. Handy if you already run your own phone setup. · Handing the call to a human with context: the AI briefs the person first, instead of a cold drop where the caller repeats themselves. · Kicking off a whole list of outbound calls at once, rather than dialling one at a time. · A standard way to let the agent use outside tools mid-call, like a booking system or your CRM. (MCP stands for Model Context Protocol.)
Speech-to-text
Speechmatics (Standard / Enhanced), Speechmatics Melia 1 (multilingual, batch-only preview)
Text-to-speech
Speechmatics TTS
Languages
en, es, fr, de, it, pt, nl, pl, ru, ar, hi, zh, ja, ko, cy
Integrations
Real-time API (streaming), Batch API (recorded files), On-prem containers (CPU / GPU), Kubernetes self-host, Virtual Appliance (on-prem VM), Native SDKs

Compliance

✓ HIPAA✓ SOC 2 Type II✓ GDPR

Our full take

Speechmatics is a speech-to-text engine first, and that is the whole point to get straight. It listens to audio and writes down the words. It now also offers its own text-to-speech, but it does not generate a reply (there is no language model) and it does not dial a phone. So if you are shopping for a finished voice agent that answers your calls, this is not that. It is one of the parts you would build that agent from, and it is a good one.

Where it earns its place is languages. Speechmatics transcribes 56+ languages off a single model, which means you get the regional accents and dialects (Brazilian Portuguese, Canadian French, and so on) without bolting on a separate pack for each. Most of the cheaper speech-to-text engines top out around seven or ten languages. Deepgram, the closest building-block vendor we cover, lists seven. If your callers speak Tagalog, Welsh, Swahili or Urdu, that gap is the entire reason to look here.

The second reason is where it runs. Most speech-to-text APIs only run in the vendor’s cloud, you send them audio and they send back text. Speechmatics will also run inside your own data centre, as a container on your own hardware (CPU or GPU), on Kubernetes, or as a pre-built virtual machine they call a Virtual Appliance. For a hospital or a bank that cannot let call audio leave the building, that on-premises option (meaning it runs on your own servers, not someone else’s cloud) is often a hard requirement, not a nice-to-have. It is the kind of thing you cannot retrofit, so it matters that it is there from the start.

Now the pricing, and here is the bit that trips people up twice over. Speechmatics bills per hour of audio, not per minute like the agent platforms. And the rate table you see is not the list price: it renders with a “Model Training” toggle switched on by default, a setting that shares your audio with Speechmatics to train its models in exchange for a discount of about a third. Every figure on the default page assumes a data deal you have not agreed to. We switched the toggle off on 23 July 2026 and read the real list rates, and those are the figures this page now stores. The list floor for the established models is $0.36 an hour on Standard, about $0.006 a minute, the figure shown at the top of this page. Treat that as the floor, not the average. The real rate climbs with the model you pick and the mode you run.

Here is how the list rates split, read with the discount toggle off. Real-time transcription (live, as the audio streams in) is $0.36 an hour on the Standard model and $0.64 on the higher-accuracy Enhanced model. Batch transcription (a recorded file after the fact) is $0.36 Standard and $0.60 Enhanced. In per-minute terms that is about $0.006 to $0.0107 depending on model and mode. Leave the model-training toggle on and the displayed rates drop to $0.24/$0.43 real-time and $0.24/$0.40 batch, the discounted figures we ourselves stored before this correction, which is rather the point about how easy the default is to mistake for the price. The rate table loads through JavaScript, so we read these figures from a dated capture rather than a plain page fetch. The free tier remains 3,000 minutes a month (50 hours), split 1,200 real-time and 1,800 batch, no card needed, plus around a million free text-to-speech characters, and a separate 20% volume discount once you cross 500 hours a month.

The new model is Melia 1, launched 2026-06-17. It transcribes multilingual audio in one pass, including speakers who switch language mid-sentence, across 55+ languages by Speechmatics’ own description, with no need to pick a language up front. Two caveats before you build on it. First, it is a batch-only production preview for now (real-time is promised, not shipped), so it cannot power a live phone agent yet. Second, the advertised as low as $0.129 an hour, which is also the headline price on the Pro plan card, is the model-training-discounted figure; the list rate with the toggle off is $0.192 an hour. Speechmatics’ launch post also claims Melia beats Deepgram, Microsoft and AssemblyAI on the FLEURS multilingual test set. That is the vendor marking its own homework, so treat it as a claim, not a result.

One honest caveat on cost. That per-minute number looks tiny next to a $0.06-a-minute agent platform, and it is, but it is not comparing like for like. Speechmatics is charging you for one job, the transcription. The platforms are charging for transcription plus the language model plus the voice plus the phone line bundled together. To build a full phone agent on Speechmatics you still have to pay for an LLM and telephony separately, though it now has its own text-to-speech ($0.011 per 1,000 characters) so the voice no longer has to come from elsewhere. Add those up and the real per-minute cost lands a lot closer to the bundled platforms than the $0.006 headline suggests.

On compliance, Speechmatics is unusually well-documented for a building block. Its own security page states SOC 2 Type II, ISO/IEC 27001:2022, GDPR and full HIPAA compliance, with AES 256 encryption at rest and TLS 1.2 or higher in transit, plus a public trust centre where you can pull the actual reports. We have ticked HIPAA, SOC 2 Type II and GDPR here because the vendor states them directly. We left SOC 2 Type I unticked: the page names Type II, not Type I, and we do not assume one from the other. For a regulated buyer, that combination of on-prem deployment plus written certifications is the strong card.

My read: Speechmatics is the one you reach for when language coverage or on-premises deployment is non-negotiable, and you have the engineering to assemble the rest of the agent around it. The voice-quality and ease-of-use scores sit lower here than for a finished platform, and that is fair, this is infrastructure, not a product you switch on. If you just want calls answered without standing up your own stack, a bundled platform will get you there faster. If you need to transcribe twenty languages, or keep the audio on your own servers, very little else competes.

The 1 to 10 scores on this page are an editorial preview, our provisional read to get the framework in place, not a measured result. We have not run Speechmatics through our own test calls yet, so there is no Voxrater latency figure here. The pricing, language, deployment and compliance detail is sourced from Speechmatics’ own pricing, security, languages and deployments pages, first captured 2026-05-31 and most recently re-captured 2026-07-23 (when we also read the rate table with the model-training discount toggled off to confirm the list rates).

Alternatives to Speechmatics

Other platforms that overlap with Speechmatics on the same kind of work, ranked by how many capabilities they share, then by cheaper all-in cost per minute. Compare any of them side by side on the compare page.

Further reading

Tracking Speechmatics? Get the next test result

We re-test and re-price the platforms we cover. Join the list and the next dated update lands in your inbox.

Newsletter launching soon.

Sources

  1. Re-captured 2026-07-30 (screenshot in evidence/): no rate change. The startup-programme banner (up to $50,000 in credits) is absent from this capture, and the pricing FAQ is retitled from 'How does Model Training work?' to 'What is the model-training discount programme?', described as 'an opt-in programme that takes 33% off Speech to Text rates'. The page carries a 'Last updated 29 July 2026' stamp. Capture caveat: the cookie banner was not dismissed in this run and overlays part of the page, and the Model Training toggle rendered ON (discounted view), so the stored list rates are still sourced to the toggle-off capture of 2026-07-23. · captured 2026-07-30
  2. Speechmatics pricing re-captured 2026-07-23 (screenshot in evidence/; the JS-rendered table again defaults the 'Model Training: enable for 33% discount' toggle ON, displaying $0.24/$0.43 real-time and $0.24/$0.40 batch). A same-day toggle-OFF re-read of the same table pixel-confirmed the LIST rates now stored: real-time $0.36/hr Standard and $0.64/hr Enhanced, batch $0.36/hr Standard and $0.60/hr Enhanced, Batch Melia 1 $0.192/hr; the TTS row ($0.011/1k chars) and the bolt-on rows (translation $0.65/hr, summaries $0.12/hr, chapters $0.40/hr, sentiment $0.12/hr, topics $0.20/hr) read the same in both toggle states. · captured 2026-07-23
  3. Speechmatics pricing re-captured 2026-07-11 (JS-rendered table read from the dated screenshot; the table renders with the 'Model Training: enable for 33% discount' toggle ON by default, so the displayed per-hour figures are the discounted ones): free tier grown to 3,000 STT min/mo (50 hrs), split 1,200 real-time + 1,800 batch, plus ~1M TTS chars; 56+ languages; Batch Melia 1 listed at $0.129/hr alongside Standard/Enhanced. · captured 2026-07-11
  4. Melia 1 launch post (published 2026-06-17): multilingual model with mid-utterance code-switching across 55+ languages, batch-only production preview ('real-time on the way'), advertised as low as $0.129/hr with 10 free hours a month; the FLEURS accuracy wins over Deepgram, Microsoft and AssemblyAI are Speechmatics' own claims, not independent results. · captured 2026-07-11
  5. Startup Program page, checked 2026-07-11: the old become-a-partner URL now redirects here; up to $50k in usage credits, cohorts capped at 20 startups, under $10M raised; no partner, reseller or affiliate content remains. · captured 2026-07-11
  6. Speechmatics pricing re-captured 2026-06-15 (rate table is JS-rendered; figures read from the dated screenshot): real-time STT $0.24/hr Standard, $0.43/hr Enhanced; batch $0.24/$0.40; free tier now 2,400 min/mo; models named Standard/Enhanced. · captured 2026-06-15
  7. Speechmatics text-to-speech (new): $0.011 per 1,000 characters, English at launch. · captured 2026-06-15
  8. Speechmatics pricing verified 2026-06-02: Pro from $0.24/hr (= $0.004/min), 2,400 free minutes/mo; speech-to-text only, so no per-minute voice-output rate. · captured 2026-06-02
  9. Speechmatics pricing page re-captured 2026-06-02 for the quarterly re-verification (screenshot in evidence/). · captured 2026-06-02
  10. Speechmatics pricing page: per-plan features (Free, Pro, Enterprise), Pro from $0.24/hr, Free 480 min/month, 6,000 hr/month cap, 24,000 hr/year enterprise discount · captured 2026-05-31
  11. Speechmatics security page: SOC 2 Type II, ISO/IEC 27001:2022, GDPR and HIPAA claims, AES 256 / TLS 1.2+, Azure + on-prem · captured 2026-05-31
  12. Speechmatics languages page: 55+ languages for speech-to-text with accent/dialect coverage · captured 2026-05-31
  13. Features and deployments: SaaS, on-prem containers (CPU/GPU), Kubernetes, Virtual Appliance, real-time + batch · captured 2026-05-31
  14. Speechmatics partner programme: build/market/sell tracks plus partner marketplace · captured 2026-05-31