By Sagar Shankaran, Founder of CallSphere
Build a 300-call test set from your own cancellations and answering-service log, grade escalation at 100%, then widen the agent one visit type at a time.
Key takeaways
A man calls the orthopedic group's main line eleven days after a total knee replacement. The office closed at five. He is not panicking — the knee is warmer than it was, he thinks he might have a temperature, and he wonders whether to keep Thursday's appointment or come in sooner. What happens in the next ninety seconds is the whole question of whether an AI phone agent belongs on that line.
If it books him for Thursday and says goodnight, you have a problem that will eventually have a name and a date attached to it. If it recognises a post-operative call inside the 90-day global period, makes no scheduling decision at all, and pages the on-call surgeon the way the answering service is supposed to, it did the job. That difference is not luck. It is whether anybody graded the thing on calls exactly like that one before it went live.
An evaluation set is a few hundred real calls pulled from your own practice's last twelve months, each one marked with the answer your best scheduler would have given, used to grade the agent before it ever speaks to a live patient. Building that set out of your own history — not out of a vendor's demo script — is the part that matured in 2026 and the part that decides how much authority the agent earns.
Owners worry about the dramatic failure. The expensive failures are duller than that, and there are two of them.
The wrong slot. A new patient with a lumbar radiculopathy referral is booked into a 15-minute established follow-up with the hand surgeon, because the caller said "back" once and "wrist" once and the person on the phone at 7 p.m. picked wrong. The front desk finds the mismatch at check-in, the patient has taken a morning off work, the surgeon loses the slot, and the scheduling supervisor spends twenty minutes rebuilding a full day. In February, with ski season running and sports slots booked three weeks out, it is not refilled.
The missed escalation. A post-operative patient describing warmth, drainage, fever or calf pain. A patient who fell. A cast that is too tight and the fingers are numb. A small fraction of the volume and all of the risk, and an answering service staffed by people who have never worked in an orthopedic office misses them more often than owners like to think.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for healthcare in your browser — 60 seconds, no signup.
Everything else — a wrong location, an injection visit booked without checking whether the patient is on a blood thinner, a workers' compensation intake taken without the claim number — sits between those two: annoying, and cheap to fix once you can see how often it happens.
Two years ago, evaluating a phone agent meant listening to a handful of calls and forming an impression. In 2026 the tooling for measuring and watching agents grew up, and it now gates a rollout rather than being something you do afterwards when a complaint arrives. You run the agent against real cases, you get a score per category, and — the genuinely new part — you can read back what it did on the ones it got wrong: what it heard, what it looked up in the schedule, which rule it followed, where it turned.
That matters for one unglamorous reason. When it fails, you can tell whether the fix is a rule change (nobody told it the spine surgeon takes no new patients on Fridays) or something you should not try to fix at all (it cannot reliably separate a post-op infection from post-op normal, so that call should never be its decision). Before, every failure looked the same from outside, and the only choices were trust it or turn it off.
flowchart TB
A["Pull 300 real calls from the last 12 months"] --> B["Mark each one: right visit type, right provider, escalate or not"]
B --> C["Run the agent against all 300 with the phone line still closed"]
C --> D{"Escalation caught 100% and visit type at or above 97%?"}
D -->|No| E["Read the step-by-step record of every miss, fix the rules"]
E --> C
D -->|Yes| F["Go live for established follow-ups and cancellations only"]
F --> G["Two weeks of daily review by the scheduling supervisor"]
G --> H["Widen to new patients with a referral already on file"]
The cases are already in the building. Do not imagine them, and do not let a vendor write them for you.
Three hundred cases is enough. Aim for roughly two hundred routine and one hundred hard on purpose — real life is heavily routine, and a test set that mirrors real life lets a dangerous agent score 96% while missing every call that mattered. Have your most experienced scheduler and the triage nurse mark the right answer for each. That session is a day of work by two people, and it is the most valuable day in the project.
Score five things, not one — and be explicit that four are trade-offs and one is not.
| Measure | Answering service, historical | Agent, first run | Required before going live |
|---|---|---|---|
| Correct visit type booked | 94% | 91% | 97% or better |
| Correct provider and location | 96% | 95% | 98% or better |
| Red-flag call escalated to a clinician | 88% | — | 100%, no exceptions |
| Booked into blocked or surgery time | 11 per month | — | 0 to 2 per month |
| Workers' comp intake with claim number captured | 71% | — | 90% or better |
Now the money, with stated assumptions. Say the agent handles 520 after-hours and overflow calls a month. At the historical 6% wrong-visit-type rate, that is 31 mis-bookings. Assume each costs 22 minutes of scheduling supervisor rework at a fully loaded $26 an hour — about $9.50 each, or $295 — and that one in four consumes a slot that never gets refilled in a busy month. Value a lost new-patient slot at $240 in collected professional revenue: 7.8 slots is $1,872. Total drag, roughly $2,167 a month. Cut the error rate to 3% and that halves. The improvement is worth about $1,080 a month, measurable again in ninety days from the same cancellation report you already run.
Post-operative calls inside the global period. Not "handled carefully" — handed straight to the on-call path. Warmth, drainage, fever, calf pain, numbness below a cast, a wound that opened. The agent's only correct behaviour is to recognise the call type and get a clinician on it.
Still reading? Stop comparing — try CallSphere live.
See the healthcare AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
Controlled substance questions. Any practice with a pain management or spine service gets refill calls at eleven at night, and those need a person and a look at the state prescription monitoring database.
Workers' compensation intake beyond the claim number and the adjuster's name. Those calls carry legal weight, the case manager on the other end is documenting the conversation, and the practice should be too.
And the angry call — surgery moved, bill wrong, imaging results never called back. The agent should recognise it, say something honest and get it to a human the same day, not solve it. Widening an agent's authority into the calls that most need a person is how a good rollout becomes a bad reputation on Google in eight weeks.
Three hundred marked cases is a workable floor for a single-specialty practice, weighted toward the hard ones. What matters more than the count is that they come from your own past twelve months, your own providers, your own visit types — an agent that scores well on someone else's cases has told you nothing about Thursday afternoon in your office.
Yes. The answering service message log, the cancellation reasons in your scheduling system and the triage nurse's callback notes get you most of the way, because they contain the outcome — the part you are grading against. Start recording going forward if your state law and posted notice allow it, and check that with counsel rather than assuming.
One hundred percent on your marked set, and even that earns only a limited rollout. Escalation is the measure you never trade against convenience or booking rate. If a vendor proposes tuning it down because the agent escalates too much, the answer is that over-escalation is a cost and a missed post-op infection is a catastrophe — not the same kind of number.
Six to eight weeks: two weeks of listen-only grading, two weeks live on established follow-ups with the scheduling supervisor reading every transcript each morning, then widening to new patients with a referral on file. Never run the widening step during your seasonal peak — not in the second week of ski season.
Run one report: every cancellation and reschedule in the last twelve months coded as wrong appointment type or wrong provider. It takes ten minutes to produce, it is the first third of your test set, and it will tell you something uncomfortable about how the current after-hours arrangement is working before any agent is involved at all.
CallSphere builds the AI voice and chat agents that answer practice phone lines and web chat, book appointments and capture new patient enquiries around the clock. The part worth insisting on with any vendor, including this one, is the sequence above: grade it on your own calls first, read back what it did on the misses, keep escalation in a clinician's hands, and widen its authority one visit type at a time.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
How an industrial equipment OEM builds a parts-desk test set from its own closed orders, grades the agent on as-built revisions, and stops wrong-part shipments.
Community management agents fail on escalations and claims, not tone. Build a 400-thread test set from your own social inbox before anything goes live.
Cattle producers can now measure a phone agent against real past calls, watch every step it took, and widen what it is allowed to do only once it earns it.
Sort six months of front-counter recordings into eight call types, write the right answer for each, then widen the agent's authority in four stages, not one.
How a design firm builds an RFI test set from closed jobs, scores citation and routing accuracy, and widens an agent's authority only when the numbers earn it.
Build a test set from your own closed files, score misses apart from false alarms, and widen an agent's authority in three gates before it drafts Schedule B.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI