What Does a Voice Agent Really Cost Per Minute?

There is no single number for what a minute costs. It comes from four separate items, and which one dominates depends on how your agent talks.

8 min readCost
A cost and performance chart displayed on a screen with several coloured series

Everyone who receives a voice agent proposal asks the same thing first: what does a minute cost? Everyone who answers gives a different number, and the reason is not that somebody is lying.

Voice agent cost per minute is not a single line item. Four separate services run at the same time behind one minute of conversation, and all four are billed independently.

The four items

Each of the four does a different job and each is metered differently, which is why a single blended figure hides more than it explains.

ItemWhat it doesHow it is billed
Speech recognitionTurns what the other side says into textAudio minutes
Language modelDecides what to answerTokens, so conversation length and history
Speech synthesisTurns the answer into audioCharacters produced or audio minutes
Phone lineActually carries the callMinutes, by direction and operator

The sum of those four is your raw cost. Platform fees, infrastructure and support sit on top of it, which is why a quoted price is only meaningful with its scope attached.

Scope is also where proposals differ most. Two vendors quoting the same figure can include entirely different items, and neither of them is being dishonest about it.

So when a provider says a minute costs a certain amount, the question that makes the number usable is which of the four items are included.

Which item dominates?

This is where it gets interesting: the dominant item is not fixed. It changes with how your agent talks, which means the same platform produces different bills for different scripts.

An agent that talks a lot

If your agent builds long sentences and explains at every turn, speech synthesis becomes the dominant item, because synthesis is billed by the volume of text produced.

The irony is worth stating plainly. A talkative agent is both more expensive and worse at selling, because a high talk ratio leaves the other side's need underexplored.

That gives you a rare alignment. Shortening the agent's turns lowers the bill and raises the discovery quality at the same time, so there is no trade-off to negotiate internally.

An agent that runs long

If calls stretch out, because the other side talks a lot and the agent waits, recognition and the phone line move ahead. Both are billed on pure duration, so even silence is billed.

The item that grows fastest

Every answer the agent gives sends the whole conversation history back to the model. As a call gets longer each new answer costs more than the last one.

That makes the model cost not linear but accelerating. The last minute of a ten minute call is noticeably more expensive than the first, and nothing on the invoice explains why.

Long calls are therefore expensive twice over: once for the duration itself and once for the growing context sent with every answer. The second cost is invisible on most dashboards.

The costs that do not appear on the quote

The four items are only the visible part. Five more are usually forgotten when a budget is put together, and together they are not small.

  • Unanswered calls: the line is billed while it rings, so an unreachable number is not free
  • Answering machines: if the agent talks to voicemail all four items run and nothing is received in return
  • Retries: how many times an unreachable number is called back multiplies the total directly
  • Analysis: summary, objection detection and scoring after the call are a separate model call
  • Storage: keeping audio and transcripts is a small but permanent line

The answering machine item deserves the most attention. A setup without voicemail detection burns part of its budget systematically, and lists with low reachability make that invisible loss larger.

The better question: cost per outcome

Comparing cost per minute is the weakest way to compare two systems, because a cheap system that produces nothing is more expensive than a costly one that books meetings.

The question worth asking is what an appointment costs you. That number is what a budget meeting can actually act on, and it reorders most vendor comparisons.

It also changes which vendor wins. A more expensive minute that books more meetings produces a lower cost per appointment, and only the second figure survives contact with a finance review.

Three figures are enough to compute it: average call duration, reachability rate, and the conversion rate from reached calls to appointments.

Price per minute is an input. Cost per appointment is an outcome. Budget meetings are about the second one.

If you do not hold those three, no per minute price tells you anything, because you cannot convert it into the only figure that matters to the business.

What to ask when you get a quote

Six questions turn a quoted rate into something comparable. Without them two proposals can differ by a factor of two while showing the same number on the front page.

  1. 1.Which of the four items are included in the quoted minute price?
  2. 2.Are unanswered calls and answering machines billed?
  3. 3.Is post-call analysis included or charged separately?
  4. 4.Is billing per second or per started minute? On short calls this difference is large
  5. 5.Does the price fall with volume, and where are the thresholds?
  6. 6.Does the provider pass through their own cost or add a fixed margin? The two carry different risks

The fourth question surprises people most often. Per started minute billing on a list of thirty second calls can double an effective rate without any line on the invoice changing.

Short calls are common in outbound work, so this rounding decision matters more here than in most telephony contracts. Ask for it in writing rather than assuming per second billing.

Where the levers actually are

Once the breakdown is visible, the levers are obvious and none of them involve renegotiating a rate card. They involve changing how the agent behaves.

  • Shorten the agent's turns: this cuts synthesis and improves discovery at the same time
  • Detect answering machines early and hang up
  • Cap retries with a written rule rather than a default
  • Set a daily ceiling so a mistake cannot run overnight

Response speed sits next to this: a faster agent holds the caller's attention and shortens the call, as covered in the article on voice agent latency.

How much of this you own directly depends on whether you build or buy, which is covered in the article on building versus buying.

And the reachability figure in the cost per appointment calculation is decided long before the call, as covered in the article on the first 24 hours.

Want to see what is inside your own calls?

Callsense makes the intent, the objection and the next step in a conversation visible. A scoping call takes 30 minutes and needs no technical preparation.

Book a scoping call