Every rate in this piece comes from the dated pricing captures on our vendor pages, checked against the live vendor pages between 30 May and 23 July 2026 (capture dates in the sources below, with screenshots in our evidence store for the July sweeps). The Speechmatics, Deepgram and Cartesia figures were re-verified 2026-07-23, including a toggle-off re-read of the Speechmatics rate table. The characters-per-minute key is our documented editorial basis, not a vendor figure, and every derived rate is labelled as derived where it appears.
Every voice AI price converts to one number: cost per finished minute of speech. The master key is that a finished minute comes out of roughly 900–1,000 characters of script, so a 10-minute job is about 9,000 characters. Convert each vendor’s unit to that and the wrappers stop mattering.
That paragraph is the whole decoder; the rest of this page applies it, one unit per section, each section ending with what the identical 10-minute job costs in that wrapper. Each unit is a currency the vendor invented; you want the exchange rate.
The master key: 900 to 1,000 characters per minute
Speech generators bill by the text you feed them, so the key conversion is characters to minutes. Our documented basis, stated on the Rime page and in the narration roundup: 1,000 characters of script is roughly 150–180 words, and a finished minute of speech comes out of roughly 900–1,000 characters. So a 10-minute script is about 9,000 characters. On a live call the agent only speaks about half the time, so a conversation minute burns roughly 450–500 characters of generation.
Two honesty notes. The key is our editorial basis, not a vendor’s, and a vendor’s own conversion can disagree (Fish Audio’s does; we work the difference through below). And it prices generation only; the tier carrying your commercial licence is often the bigger line.
Credits: ElevenLabs, Cartesia, Fish Audio
A credit is a prepaid balance you spend as you generate, and it means whatever the vendor says it means.
ElevenLabs keeps it clean: on the standard models one character costs one credit, and Flash bills at half the per-character rate by its own docs. The API prices the same units in money: $0.10 per 1,000 characters on Multilingual, $0.05 on Flash. On subscription the effective rate depends on the tier, and only holds if you use the whole allowance:
| ElevenLabs tier | Price per month | Credits included | Effective rate per 1,000 |
|---|---|---|---|
| Starter | $6 | 30,000 | $0.20 |
| Creator | $22 | 121,000 | about $0.18 |
| Pro | $99 | 600,000 | about $0.165 |
Cartesia meters by time instead: about 750 credits per generated minute of speech, no flat per-character rate any more (the 15-credits-a-second line on its page belongs to the voice changer, not TTS). Pro is $5 a month for 100,000 credits, Startup $49 for 1.25 million, Scale $299 for 8 million.
Fish Audio’s subscriptions are credits too (Plus $15 for 250,000, Pro $100 for 2 million), and its plan page cannot agree with itself on what one buys: the FAQ says a minute costs 600–625 credits, while the plan cards imply roughly 1,250 (4,000 on the top tier). Same page, same capture, 11 July 2026.
The 10-minute job, in credits. ElevenLabs: 9,000 characters = 9,000 credits, so at the API’s $0.10 per 1,000 that is $0.90 (Flash: $0.45). Cartesia: 10 minutes × 750 = 7,500 credits, costing $0.38 on Pro ($5 ÷ 100,000 × 7,500), about $0.29 on Startup and $0.28 on Scale (our stored derived $0.035 per 1,000 sits inside that band; Pro runs nearer $0.04). Fish on its FAQ’s conversion: 6,000–6,250 credits, about $0.36–0.38 on Plus; on its cards’ implied conversion: 12,500 credits, about $0.75. One page, one job, double the price.
Characters, billed as characters
Speechify is the simple case: Starter at $10 a month includes a million characters, then $10 per million after, which is $0.01 per 1,000; the overage rate falls to $8 per million on Pro and $6 on Scale. Hume meters its Octave engine at $0.15 per 1,000 characters on the entry tiers, falling to $0.05 on the $500-a-month Business plan. Rime publishes a single “starting at” rate, cut from $0.05 to $0.03 per 1,000 characters in late July 2026, with volume discounts it no longer spells out. And Telnyx’s cheapest listed voice is $0.000009 per character, which looks alarming and is $0.009 per 1,000.
The 10-minute job, in characters. Speechify: 9 × $0.01 = $0.09 (about $0.054 at Scale overage). Hume: 9 × $0.15 = $1.35 at entry ($0.45 on Business). Rime: 9 × $0.03 = $0.27, a starting-at figure (it was 9 × $0.05 = $0.45 until late July). Telnyx floor: 9,000 × $0.000009 = about $0.08; premium voices cost more per character.
Bytes: Fish Audio’s API
Fish’s API does not bill characters at all: it bills $15.00 per million UTF-8 bytes on every current model. A byte is the smallest unit a computer stores, and in the UTF-8 text encoding an English letter or space takes exactly one byte, so for English scripts bytes and characters are the same number. Other scripts are not so lucky: under the encoding standard (RFC 3629), characters outside the basic Latin set take 2–4 bytes each, so a Chinese or Japanese script at three bytes per character costs roughly three times the English figure. That is the encoding standard, not a Fish surcharge, but it lands on the bill.
The 10-minute job, in bytes, twice. By our key: 9,000 English characters is about 9,000 bytes, so 9,000 ÷ 1,000,000 × $15 = about $0.14. By Fish’s own conversion (“1M UTF-8 bytes is approximately 180,000 English words, or about 12 hours of speech”): $15 ÷ 720 minutes ≈ $0.021 per minute, so 10 minutes = about $0.21. The two disagree because Fish’s conversion implies about 1,389 bytes per spoken minute (1,000,000 ÷ 720) against our 900–1,000; neither is wrong, they just assume different speech rates. Budget on the $0.21 and be pleased if your script comes in nearer $0.14.
Bundled hours: Murf
Murf sells finished audio by the year. Creator is $19 a month billed annually, which is $228 a year for 24 hours of generation; Business is $66 a month, $792 a year, for 96 hours. Murf publishes no per-character or per-minute rate, so the division is ours: $228 ÷ 1,440 minutes ≈ $0.16 a minute on Creator, $792 ÷ 5,760 ≈ $0.14 on Business. The shape matters as much as the rate: bundled hours are pre-bought, so an allowance you do not use is money already spent.
The 10-minute job, in bundled hours. Creator: 10 × $0.158 ≈ $1.58. Business: 10 × $0.1375 ≈ $1.38. Our narration roundup prints $1.44 for the same job because it derives per 1,000 characters ($0.16 × 9) rather than per bundled minute. Both are honest divisions of the same $228; the gap is rounding inside the characters-per-minute key, so we show both rather than average them.
Per minute, with and without passthrough: live calls
Live call platforms quote per minute, and the question is what the minute includes. Vapi charges a $0.05-a-minute platform fee and passes the rest through: speech-to-text, the model, the voice and the phone line billed at cost from the providers you wire in (passthrough means no markup on those parts). Our recorded all-in band for a realistic stacked build is $0.05–0.30 a minute. Retell’s own banner says $0.07–0.31 all-in, and its published floor adds up: $0.055 infrastructure + $0.015 voice + $0.003 cheapest model + free SIP ≈ $0.073. Speechify’s agents go the other way: from $0.07 a minute all-in, no passthrough, no token maths, says the pricing page.
One warning: a call minute is not a narration minute. The price covers listening, reasoning, speaking and the phone line while both sides talk, so never compare it against a per-character rate. We dissect the call minute in what a voice agent really costs per minute.
The 10-minute job, as a call. Vapi: 10 × $0.05–0.30 = $0.50–3.00, your stack decides. Retell: 10 × $0.07–0.31 = $0.70–3.10 (a typical GPT 4.1 and Twilio build runs about $0.13 a minute, so $1.30). Speechify agents: 10 × $0.07 = $0.70 flat.
Tokens: OpenAI’s Realtime API
A token is the small chunk of data an AI model reads and writes, and OpenAI Realtime bills voice in them. The conversion is vendor-documented rather than folklore: OpenAI’s own cost guide states that input audio costs 1 token per 100ms and output audio 1 token per 50ms, so 600 input tokens and 1,200 output tokens per spoken minute. The audio rates are $32 per million input tokens ($0.40 cached) and $64 per million output tokens.
The 10-minute job, in tokens (a call, half caller and half agent). Five minutes of caller audio in: 5 × 600 = 3,000 tokens, × $32 per million ≈ $0.10. Five minutes of agent audio out: 5 × 1,200 = 6,000 tokens, × $64 per million ≈ $0.38. Raw audio total: about $0.48. And honestly, the floor is not the bill: your instructions plus the conversation so far are re-sent as text context on every turn, and that usually dominates. One independent analysis we track puts the realistic range at $0.18–0.46 a minute uncached and $0.05–0.10 with prompt caching, so $1.80–4.60 or $0.50–1.00 for the call; the analyst’s model, not an OpenAI rate.
Hours of audio: the transcription side
Speech-to-text vendors price the listening, usually per hour of audio, and the model labels matter. AssemblyAI’s real-time streaming starts at $0.15 an hour on Universal-Streaming; its asynchronous transcription (after-the-fact, not live) runs $0.15 an hour on Universal-2 and $0.21 on Universal-3.5 Pro; and the premium real-time model, Universal-3.5 Pro Realtime, is $0.45 an hour. Quote “$0.45 streaming” unqualified and you overstate the entry price three times. Deepgram meters the same job per minute instead, and Speechmatics shows discounted rates by default (more in the traps below).
| Vendor | Unit | Base real-time rate | The 10-minute job |
|---|---|---|---|
| AssemblyAI | per hour | $0.15/hr ($0.0025/min) | $0.025 (premium model: $0.075) |
| Speechmatics | per hour | $0.36/hr list ($0.24 with the data-sharing discount ON) | $0.06 ($0.04 discounted) |
| Deepgram | per minute | $0.0077/min Nova-3 streaming list ($0.0048 limited-time promo) | $0.077 ($0.048 promo) |
The 10-minute job, transcribed. 10 × $0.0025 = $0.025 on AssemblyAI’s base streaming ($0.075 premium); 10 × $0.0077 = $0.077 on Deepgram Nova-3 at list ($0.048 on the limited-time promo). Listen-only costs: a whole agent also reasons and speaks, so a $0.006 transcription minute and a $0.07 agent minute are not rivals.
The master conversion table
The same 10-minute job in every unit, each figure worked in its section. The call and transcription rows price a different job from the narration rows: a unit decoder, not a ranking.
| Unit | Who uses it | What the 10-minute job costs |
|---|---|---|
| Credits | ElevenLabs (1 character = 1 credit), Cartesia (about 750/minute), Fish Audio plans | $0.90 ElevenLabs standard ($0.45 Flash); $0.28–0.38 Cartesia by tier; $0.36–0.75 Fish, conversion-dependent |
| Characters | Speechify, Hume, Rime, Telnyx | $0.09 Speechify; about $0.08 Telnyx floor; $0.45 Rime starting-at; $1.35 Hume entry |
| UTF-8 bytes | Fish Audio’s API | about $0.14 by our key; about $0.21 by Fish’s own conversion |
| Bundled hours | Murf studio plans | about $1.58 Creator, $1.38 Business (our division) |
| Per minute (call) | Vapi, Retell, Speechify agents | $0.50–3.00 Vapi; $0.70–3.10 Retell; $0.70 Speechify |
| Tokens | OpenAI Realtime | about $0.48 raw audio; $0.50–4.60 realistic with text context |
| Hours of audio (STT) | AssemblyAI, Speechmatics; Deepgram per minute | $0.025–0.075 AssemblyAI; $0.06 Speechmatics list ($0.04 discounted); $0.077 Deepgram list ($0.048 promo); listen-only |
The traps our captures keep catching
Six patterns, all caught on live pricing pages, all with dated screenshots.
The page that disagrees with itself. Fish’s FAQ says 600–625 credits a minute; its own plan cards imply roughly 1,250. When a page self-contradicts, price on the dearer reading.
The default-ON discount toggle. Speechmatics’ rate table renders with a “Model Training” toggle enabled, a 33 per cent discount in exchange for your audio training its models. The displayed price assumes a data deal you have not agreed to. Toggle it off before quoting: at our 23 July 2026 re-read, off meant $0.36 an hour for real-time Standard, not the $0.24 the table leads with. Deepgram now runs the same pattern twice over, a limited-time promo price with the list struck through, on rates that also assume its Model Improvement Program data-sharing opt-in.
The promo price where the regular price belongs. ElevenLabs’ Creator card showed $11 at our 11 July capture, a first-month offer; the regular price is $22. We briefly recorded the promo as the base ourselves, which is rather the point: it catches people who compare prices for a living.
Annual pricing dressed as monthly. Murf’s “$19 a month” Creator plan bills annually, a $228 commitment up front. Cartesia and Fish have shown annual-equivalent monthly figures too. Check for “billed annually” near the big number.
The geo-priced page. Fish’s plan page rendered in pounds from a UK connection until we switched its currency picker to USD. Compare it against a dollar page unaware and you are off by the whole exchange rate.
The starting-at rate. Rime’s “STARTING AT” rate comes with unpublished volume discounts, and it moves: $0.05 per 1,000 characters at our 23 July capture, $0.03 a week later, with the FAQ on the same page still quoting the old number. A fine floor, a poor budget. Get the rate card in writing.
How to compare any two platforms in three steps
Step one: find the real unit price. Not the plan price, the per-unit price: dollars per 1,000 characters, per million bytes, per minute or per hour. If the vendor only sells bundles, divide the price by the units inside and call the result your own division, as we do with Murf.
Step two: convert to cost per finished minute. Use the key: 900–1,000 characters per finished minute, about half that per conversation minute for a live agent. If the vendor publishes its own conversion, run both, as we did for Fish; when they disagree, budget on the dearer one.
Step three: price your real month, not the worked example. Multiply by your actual volume, then add the subscription floor and the allowance you will not use. This is where bundled hours and use-it-or-lose-it credits quietly change the answer.
Our calculator runs all three steps against every platform we track, using the stored, dated rates behind this page. Put your own volume in and let it argue with your shortlist.
Common questions
How do ElevenLabs credits work?
How many characters is one minute of AI speech?
Why do prices on the same vendor page disagree with each other?
Is a per-minute call rate comparable with a per-character narration rate?
Sources
Every figure above is dated and links to its primary source.
- ElevenLabs pricing page (captured 2026-07-11, screenshot in evidence/): 1 character = 1 credit on the standard models; tier allowances Free 10,000 credits, Starter $6/30,000, Creator $22/121,000, Pro $99/600,000, Scale $299/1.8M, Business $990/6M; the Creator card displayed an $11 first-month promo against the $22 regular price. checked 2026-07-11
- ElevenLabs API pricing page (fresh capture 2026-07-12): Multilingual v2/v3 at $0.10 per 1,000 characters, Flash at $0.05 per 1,000. checked 2026-07-12
- ElevenLabs models docs (fresh capture 2026-07-12): Flash billed at a 50% lower price per character for API generations. checked 2026-07-12
- Cartesia pricing page (captured 2026-07-11, re-read 2026-07-23, screenshot in evidence/): TTS metered in credits with no flat per-character rate published; the page's '15 credits per second' line is the voice-changer rate, and the plan cards' own maths implies about 750 credits per generated minute of TTS (100K credits ≈ 133 min on Pro, 1.25M ≈ 1,667 on Startup, 8M ≈ 10,667 on Scale); Pro $5/mo with 100,000 credits, Startup $49 with 1.25M, Scale $299 with 8M. checked 2026-07-11
- Fish Audio plan page (captured 2026-07-11 after switching its currency picker from geo-rendered GBP to USD): subscription credits Free 8,000/mo, Plus $15/250,000, Pro $100/2M, Max $999/25M; the FAQ states 600 to 625 credits per minute while the plan cards imply about 1,250 (Max: 4,000), on the same page. checked 2026-07-11
- Fish Audio API pricing docs (re-confirmed 2026-07-12): $15.00 per 1M UTF-8 bytes on every current TTS model, with the vendor's own conversion '1M UTF-8 bytes is approximately 180,000 English words, or about 12 hours of speech'; Fish STT at $0.36 per audio hour. checked 2026-07-11
- RFC 3629, the UTF-8 encoding standard (captured 2026-07-12): ASCII characters encode as 1 byte; characters outside the basic Latin range take 2 to 4 bytes each. checked 2026-07-12
- SpeechifyAI pricing page (captured 2026-07-11, screenshot in evidence/): Starter $10/mo including 1M characters then $10 per 1M, Pro $99 including 3M then $8 per 1M, Scale $499 including 10M then $6 per 1M; Voice Agents from $0.07/min all-in with tier overage down to $0.068 and Enterprise from $0.06. checked 2026-07-11
- Hume pricing page (re-captured 2026-06-15): Octave TTS at $0.15 per 1,000 characters on the entry tiers, down to $0.05 per 1,000 on the $500/mo Business plan. checked 2026-06-15
- Rime pricing page (captured 2026-07-11, screenshot in evidence/): a single 'STARTING AT $0.05/1K characters' rate with volume discounts unpublished; the earlier per-model rates (Mist at $0.03) removed in the mid-2026 redesign, with placeholder FAQ text still on the page. checked 2026-07-11
- Rime pricing page re-captured 2026-07-30 (screenshot in evidence/): the Starter card now reads 'STARTING AT $0.03 / 1K CHARACTERS', cut from $0.05 at our 2026-07-23 capture, which puts the single published rate back at the old per-model Mist level. The FAQ on the same page still quotes $0.05, so the page contradicts itself; we record the card and flag the conflict. checked 2026-07-30
- Telnyx voice AI page (captured 2026-07-11): cheapest listed per-character voice (Amazon Polly standard) at $0.000009 per character, $9 per million. checked 2026-07-11
- Murf pricing page (captured 2026-05-30, tiers re-checked 2026-06-15): Creator $19/mo billed annually ($228/yr) for 24 hours of generation a year; Business $66/mo ($792/yr) for 96 hours a year; no flat per-minute or per-character rate published. checked 2026-05-30
- Vapi pricing page (re-verified 2026-06-15, screenshot in evidence/): $0.05/min platform fee with speech-to-text, the model, the voice and telephony passed through at cost; our stored all-in band is $0.05 to $0.30/min. checked 2026-06-15
- Retell pricing page (captured 2026-07-11, screenshot in evidence/): banner all-in $0.07 to $0.31/min; published components $0.055/min voice infrastructure with speech-to-text included, text-to-speech $0.015/min, GPT 5 nano $0.003/min, SIP free. checked 2026-07-11
- OpenAI API pricing (captured 2026-07-11): gpt-realtime audio tokens at $32.00 per 1M input, $0.40 per 1M cached input, $64.00 per 1M output. checked 2026-07-11
- OpenAI realtime cost guide (fresh capture 2026-07-12): input audio costs 1 token per 100ms and output audio 1 token per 50ms, which is 600 input tokens and 1,200 output tokens per spoken minute. checked 2026-07-12
- AssemblyAI pricing page (re-verified 2026-07-11): Universal-Streaming real-time at $0.15/hr; asynchronous Universal-2 at $0.15/hr and Universal-3.5 Pro at $0.21/hr; the premium Universal-3.5 Pro Realtime model at $0.45/hr. checked 2026-07-11
- Deepgram pricing page (re-verified 2026-07-11): speech-to-text metered per minute, Nova-3 streaming at $0.0048/min. checked 2026-07-11
- Deepgram pricing page (re-captured 2026-07-23, screenshot in evidence/): the streaming table now labels its rates 'limited-time promotional rates on streaming' with the list price struck through, Nova-3 mono $0.0077/min list against the $0.0048 promo; a footnote states the listed rates opt in to the Model Improvement Program (data-sharing). checked 2026-07-23
- Speechmatics pricing page (captured 2026-07-11, JS-rendered table read from the dated screenshot in evidence/): real-time Standard at $0.24/hr, displayed with the 'Model Training' 33% discount toggle ON by default, so the shown rates are the discounted ones. checked 2026-07-11
- Speechmatics pricing page (re-captured 2026-07-23 with a same-day toggle-OFF re-read of the rate table): list rates pixel-confirmed at $0.36/hr real-time Standard and $0.64/hr Enhanced against the $0.24/$0.43 the default toggle-ON view displays. checked 2026-07-23
- Independent (secondary) cost analysis of OpenAI Realtime across 11 call profiles: a realistic $0.18 to $0.46/min uncached and $0.05 to $0.10/min with prompt caching. An analyst's model, not an OpenAI rate. checked 2026-05-31
Get the next piece
New analysis and dated test results land in the newsletter first. No spam.
Newsletter launching soon.