Metilda Clinical Insights
Why AI Receptionists Sound Robotic to Allied Health Patients (And How to Test for It)
An AI receptionist sounds robotic for one specific, measurable reason, not a vague one: the pause between when a caller stops talking and the AI starts responding is too long for the human brain to read as a real conversation. Script quality, voice cloning and a warm tone don't fix that if the gap itself is wrong, and most allied health clinics evaluating an AI receptionist have never actually timed that gap on a real call.
That gap has a name in the industry that builds these systems: response latency. On 28 September 2026, one of the largest voice-AI vendors published a new benchmark for one piece of that gap, and it's a genuinely large jump. It is also, on its own, a much smaller fix than the headline makes it sound, which is worth unpacking before any clinic assumes it solves the problem.
Why AI Receptionists Sound Robotic: It Is Not the Script, It Is the Silence
The same patient, the same sentence, two different outcomes, decided entirely by a number most allied health clinics never ask a vendor for: how many milliseconds of silence sit between the two.
Human conversation has a rhythm most people have never consciously noticed, because it's automatic. Cross-linguistic research on conversational turn-taking, published in the Proceedings of the National Academy of Sciences, found that the gap between one person finishing a turn and the next person starting one averages around 200 milliseconds across the languages studied, close to the floor of what a human can physically plan and execute a response in. That's the benchmark a phone-based voice AI is quietly competing against, whether or not anyone building it says so out loud.
The telecom industry has its own version of this threshold. Parloa's breakdown of voice AI latency for contact centres cites the ITU-T's G.114 standard, which holds that one-way delay under roughly 150 milliseconds is generally sufficient for "transparent" conversational interactivity, something like 300 milliseconds round trip before a delay becomes perceptible. A related standard, G.1051, puts the point where two-way delay makes an interaction "very difficult" at around 250 milliseconds. Parloa lines those telecom thresholds up against where the voice AI industry actually sits today: a median response time of 1,400 to 1,700 milliseconds. That's five to eight times slower than the natural human gap, and it's a large part of why a caller interrupts, repeats themselves, or says "hello? hello?" into a phone that's still thinking.
A clinic doesn't need to know any of those numbers to feel the effect. A patient calling after hours about a sore shoulder experiences it as a beat of dead air that makes them wonder if the call dropped, then talk over the reply when it finally starts, then decide the whole exchange felt like talking to a machine. At that latency, it did.
The Latency Budget: Where Those Milliseconds Actually Go
A phone call with a voice AI isn't one measurement, it's a chain of them, and that matters because fixing one link in the chain doesn't automatically fix the others. A latency breakdown published by WebRTC Ventures on 23 September 2026 walks through where the time actually goes, stage by stage, landing on roughly 800 milliseconds as the rough budget before a conversation starts to feel slow to the person on the other end.
WebRTC Ventures' breakdown names speech-to-text and the language model's time-to-first-token as the two biggest contributors to total round-trip latency (shaded orange above), not the network and not the final text-to-speech step, which is the part most AI receptionist marketing actually talks about.
That last point is the one worth sitting with. A faster, more natural-sounding voice model only speeds up one stage of a six-stage chain. If the turn-detection is slow to decide a caller has finished talking, or the transcript takes a beat too long to produce, or the language model is still composing its reply, a better voice at the end of that chain still arrives late. It'll sound more pleasant when it gets there. It won't feel faster.
| Benchmark | Figure | What it represents |
|---|---|---|
| Natural human conversation gap | ~200ms | Average cross-linguistic turn-taking gap (PNAS) |
| ITU-T G.114 "transparent" threshold | ~150ms one-way (~300ms round trip) | Telecom standard for imperceptible delay |
| ITU-T G.1051 "very difficult" threshold | ~250ms two-way | Point at which interaction becomes noticeably hard |
| Industry median voice AI response (2026) | 1,400–1,700ms | Parloa's cited industry benchmark |
| Eleven v4 Turbo median inference latency | ~100ms | ElevenLabs' own published figure, one stage of the chain |
What Changed on 28 September 2026
ElevenLabs announced Eleven v4 Turbo three days before this was written, pitching it specifically at "a voice that has to answer while someone is waiting, such as a phone agent or a live assistant." The company's own published figure is a median inference latency of around 100 milliseconds for the text-to-speech stage, with the model ranked first by independent benchmarking firm Artificial Analysis and preferred by roughly 75% of listeners in blind head-to-head tests against competing models including Cartesia Sonic 3.6 and two Google Gemini TTS variants. It supports more than 90 languages.
Worth saying plainly: that benchmark comes from ElevenLabs itself, blind listener panel or not, so it's a vendor marking its own homework. But the figure that doesn't need to be taken on faith is the one in the latency budget above. If the text-to-speech stage alone can now run at roughly 100 milliseconds instead of being a meaningful fraction of a 1,400-millisecond total, that stage stops being where the problem usually lives. The two stages that remain the biggest contributors, per WebRTC Ventures, are speech-to-text and the language model's own thinking time, and neither of those improved on 28 September. A clinic shopping for an AI receptionist on the strength of "our voice model just got faster" is hearing about one-sixth of the chain.
Marketing Claims vs. Measurable Latency
Cliniko's own Artificial Intelligence connected-apps directory currently lists 35 apps, 13 of which market themselves specifically as AI receptionist or call-answering products for clinics running Cliniko. Reading through those 13 listings, almost every one claims some version of sounding natural or human: "like a real human," premium, empathetic, warm. Exactly one, Intavia, states the specific claim "natural-sounding voice" in its listing copy. None of the 13 publish a response-time figure, a latency benchmark, or anything a clinic could check without picking up the phone.
That's not a knock on any one of them specifically. It's a structural fact about how this category markets itself. "Sounds human" is a subjective claim a vendor writes about their own product. Response latency is a number a clinic can measure on a two-minute phone call, with a vendor's own live demo line and a stopwatch, before signing anything.
How to Test an AI Receptionist's Latency on a Live Demo Call
A two-minute test any clinic can run before signing a contract, using nothing but the vendor's own published demo line and a phone's stopwatch.
A practice manager doesn't need engineering background to run this test. Call the number, say something a real patient would actually say, like "hi, I need to reschedule my Thursday appointment," then stop talking cleanly, mid-breath, the way people do on the phone. Start timing from the moment you stop. A reply that starts within roughly a second, without an awkward false start or the AI talking over the tail end of your sentence, is in the range this post's benchmark table calls natural. A reply that makes you wonder if the call dropped, or that starts and then restarts, is sitting closer to that 1,400–1,700 millisecond industry median than to anything resembling a real conversation.
Two follow-up questions are worth asking directly, because the answers reveal more than the marketing copy does. First: interrupt the AI mid-sentence on a second test call and see whether it yields gracefully or keeps talking over you. That's a turn-detection problem, not a voice problem, and no amount of a nicer-sounding TTS model fixes it. Second: ask the vendor what speech-to-text and language model they're running, and whether that infrastructure is hosted in a region close to Australia. None of this requires taking anyone's word for anything. It's a phone call and a stopwatch.
Frequently Asked Questions
Why does my AI receptionist sound like it's talking over patients or pausing awkwardly?
That's almost always a latency problem, not a script or voice-quality problem. If the gap between a caller finishing a sentence and the AI replying runs anywhere near the 1,400–1,700 millisecond industry median, the caller's brain reads the silence as "did it hang up" and starts talking again right as the AI starts replying, which produces the talk-over effect. Fixing the voice model doesn't fix this if the delay is happening earlier in the chain, at turn-detection, transcription or the language model's own response time.
What response time counts as "natural" for an AI phone receptionist?
Human conversation runs on a gap of roughly 200 milliseconds between turns, and telecom engineering standards (ITU-T G.114) treat delay under about 150 milliseconds one-way, or 300 milliseconds round trip, as imperceptible. In practice, a response that begins within about a second of the caller finishing, with no false start, reads as natural to most callers. That's well short of perfect, but well clear of the 1,400-plus millisecond median that makes an AI receptionist sound like a bot.
Does a faster text-to-speech model fix all of an AI receptionist's latency problems?
No, and that's the part vendor announcements tend to leave out. Text-to-speech is one stage in a chain that also includes turn-detection, speech-to-text, language-model response time and network transport. Independent analysis from WebRTC Ventures names speech-to-text and the language model's own thinking time as the two biggest contributors to total latency, not the voice stage. A faster voice model helps, but a clinic should still ask what the rest of the chain looks like.
How can I test an AI receptionist's latency before signing a contract?
Call the vendor's own live demo line, say a normal sentence a patient would say, stop talking cleanly, and time the gap with a phone stopwatch until the reply begins. A gap under about a second with no awkward false start is a good sign. Also try interrupting it mid-reply on a second call to see whether it yields naturally; that tests turn-detection, which a fast voice model alone won't fix.
Is Metilda built on this kind of low-latency approach?
Metilda, Langoedge's AI receptionist built specifically for Cliniko clinics, ships with a public live voice demo a clinic can call and test directly, rather than asking a practice manager to take a marketing claim on faith. The deeper part of Metilda's "humane" positioning isn't only about how fast the first word of a reply arrives. It's what happens next in the call, including retrying a booking against the live Cliniko calendar instead of dead-ending into a scripted failure when a requested slot is already gone. Latency decides whether the first few seconds of a call feel human. What the AI does after that decides whether the rest of the call does too.
Pick Up the Phone and Time It Yourself
Every AI receptionist vendor in Cliniko's own directory describes itself as natural-sounding, human-like, or warm, and not one of the 13 receptionist-specific listings backs that up with a number a clinic can check. That gap between the marketing copy and the measurable claim is the whole story here. The fix isn't reading a comparison article or taking a sales call's word for it — it's calling the vendor's own demo line, saying one ordinary sentence, and watching the clock. A clinic that does that for every product on its shortlist will learn more in ten minutes than any feature table on this topic could tell it.
Sources
- ElevenLabs — "Eleven v4 Turbo is now available in ElevenAgents" (28 September 2026)
- Parloa — "Speech latency in voice AI for CX"
- WebRTC Ventures — "The Voice AI Latency Budget: Where Every Millisecond Goes" (23 September 2026)
- Cliniko Connected Apps — Artificial Intelligence category
- Proceedings of the National Academy of Sciences — "Universals and cultural variation in turn-taking in conversation" (Stivers et al., 2009)