By Sagar Shankaran, Founder of CallSphere
Before an AI answers your PT clinic's intake line, build a 150-call answer key from last year's schedule. What to sample, how to score, where humans stay.
Key takeaways
It is 7:52 on a Monday in October at a four-therapist clinic and three lines are lit. Line one is a Medicare patient moving her Thursday to Friday. Line two is a self-pay runner asking what an evaluation costs and whether she needs a doctor first. Line three is the orthopedic office: a post-op ACL reconstruction is being discharged today, needs to start Wednesday, and the script will be faxed "sometime this afternoon." The patient access coordinator has one headset, a filling waiting room, and a therapist at the desk asking whether the 8:00 eval showed.
Every AI voice vendor points at this hour, and they are right about the hour. They are usually wrong about what happens next. Whether an intake agent helps your clinic or quietly bleeds cases out the back has little to do with how natural it sounds. It has everything to do with whether it knows that the ACL on line three cannot be treated under Medicare Part B without a certifying physician signature on the plan of care inside 30 days, and that the runner on line two can be seen under your state's direct access rule while a Medicare beneficiary cannot. None of that shows up in a demo. All of it shows up in your denial report ninety days later.
Through 2024 and 2025, buying an AI phone agent meant sitting through a demo, liking it, switching it on, and finding out. The tooling that grew up in 2026 changed the order of those steps. Evaluation and observability — scoring an agent against a fixed set of real past cases, then watching each live conversation step by step to see what it looked up and what it decided before it spoke — became the thing that gates a rollout rather than a nicety bolted on afterward. You measure, you watch, and only then do you widen what the agent may do without a person in the loop.
An evaluation set — in plain English, an answer key — is a fixed list of real past cases from your own clinic where you already know the correct outcome, used to grade an AI agent before it is ever allowed to speak to a patient.
The good news for a rehab practice is that you are already sitting on the raw material. Your history in WebPT, Prompt, Raintree, Jane, TheraOffice or Net Health records what actually happened to every intake call last year: the payer, whether an authorization was required and obtained, how many visits were approved, whether a referral was on file, which clinician the patient landed with, and whether the case survived past visit three. Nobody has to invent test cases. You have twelve months of them, with the right answers already written down by the people who did the work.
Pull a stratified sample, not a random one. A random 150 calls from a busy clinic is ninety percent easy reschedules and teaches you nothing. You want the mix that actually breaks front desks:
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for healthcare in your browser — 60 seconds, no signup.
For each, the office manager and the billing specialist write the correct outcome in a spreadsheet: what the agent should have said, what it should have booked, what it should have refused to answer. That spreadsheet is the deliverable. Two people, about six hours across a week, and it outlives any vendor you sign with.
flowchart TD
A["Pull 150 real intake calls from last year's schedule"] --> B["Office manager and biller write the correct answer for each"]
B --> C["Agent runs all 150 with the phone line still off"]
C --> D{"Payer, referral rule and eval slot all correct?"}
D -->|Below 95 percent| E["Read the step-by-step record, fix the rules, rerun"]
E --> C
D -->|95 percent or better| F["Answer overflow only, 7 to 9 a.m."]
F --> G["Clinic director reviews 20 live calls every Friday"]
G --> D
Do not score the agent on a single percentage. A rehab intake call has four separable decisions with wildly different consequences. Grade each separately.
Identification — did it get the payer, the body part and the referring provider right? A miss here is usually recoverable at the desk. Eligibility — did it correctly decide whether a referral or a certification was needed before the first visit? A miss here produces a visit you cannot bill. Booking — did it put the patient with a clinician who can legally and clinically treat them, at a time your schedule supports? A miss here burns an evaluation slot, the scarcest thing most clinics own. Restraint — did it decline to answer things it should not? An agent that confidently quotes a remaining deductible from stale eligibility data costs more in goodwill than it ever saved in labor.
The step-by-step record is what makes this fixable rather than mysterious. When the agent tells a Medicare Advantage patient she is approved for twelve visits, you can open that conversation and see which eligibility check it ran and where it drew the wrong conclusion. In 2024 you got a transcript and a shrug.
Why the answer key pays for itself, in round numbers you should replace with your own. All figures below are illustrative assumptions, not measured results.
| Assumption | Value |
| New evaluations booked per month | 60 |
| Average visits per completed case | 9 |
| Average net collection per visit | $95 |
| Revenue per completed case | $855 |
| Intake calls the agent handles unsupervised | 60% |
| Bad-booking rate before evaluation and tuning | 4% |
| Bad-booking rate after tuning against 150 cases | 1% |
| Share of bad bookings that never rebook | 50% |
Sixty evals at sixty percent is 36 calls a month the agent owns. At four percent error that is 1.44 bad bookings; at one percent, 0.36. The difference is roughly one a month, half of which walk permanently: about 0.54 lost cases, or $462 a month, $5,540 a year — before the denied visits you write off and the therapist hour that sat empty. Against six hours of office-manager time and a few hours a month of call review, the math is not close, and the answer key gets reused every time a payer changes its rules.
An evaluation set only covers what already happened. It cannot grade the agent on a payer policy that changes next Tuesday, or a patient describing symptoms that belong in an emergency department rather than on your Wednesday schedule. Three things stay with people no matter how good your scores get.
Still reading? Stop comparing — try CallSphere live.
See the healthcare AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
First, anything that sounds like a red flag. Unexplained calf swelling after a knee replacement, new bowel or bladder changes with low back pain, chest pressure on exertion — the agent's only correct move is to stop, say it cannot advise, and transfer to a licensed clinician or tell the caller to seek immediate care. Score restraint hard on those fifteen ugly calls.
Second, anything that commits money. Quoting a remaining deductible in January, when half your patients have reset and your eligibility data is a week stale, is how a clinic ends up eating balances. Third, the discharge and plan-of-care conversation. Recertification at 90 days, the progress note at the tenth visit, whether a plateaued patient should be discharged — clinical judgments belonging to the treating physical therapist, as your state practice act says plainly.
A working number, not a rule. What matters is that every category has 15 to 20 examples; below that, one wrong answer swings the score more than the truth warrants. A single-site clinic with two payers gets away with 80. A three-site practice taking comp, no-fault and two Medicare Advantage plans wants closer to 250.
You can use your own recordings and scheduling history for your own operations, but a vendor doing the scoring is handling protected health information, so you need a signed business associate agreement before anything leaves your building. Many practices sidestep it by de-identifying the answer key — case numbers instead of names and dates of birth — since the agent is graded on the decision, not the identity.
Set the bar by consequence, not by an average. Most clinics land near 98 percent on eligibility and restraint, 95 percent on booking, and whatever you get on identification, since the desk catches those. If eligibility sits at 90 percent you are not ready, no matter how the rest scores.
Do not shop for a vendor first. Open your scheduling system, export every new-patient intake from the same month last year, and sort it by payer. Within an hour you will know whether your hard cases are workers' comp, Medicare Advantage, or the physician offices that fax scripts late. That distribution is the first column of your answer key.
CallSphere builds AI voice and chat agents that answer clinic phone lines and web chat, book appointments and capture new-patient inquiries around the clock — including the 7:00 a.m. overflow and the Saturday calls that currently reach voicemail. If you are evaluating any agent for your intake line, ours included, build the answer key first and make the vendor score against it. A vendor who will not be measured against your own history is telling you something.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Knoxville behavioral-health clinics face deep intake backlogs. AI phone triage answers every new-patient call, screens fit, and books evaluations 24/7.
The EU AI Act's August 2 date, Texas TRAIGA and California SB 53 all landed. What a US rehab clinic must document, disclose and log, and what it can skip.
Seven in ten PT referrals are routine. Route comp, PIP and delegated Medicare Advantage to the strong model, and let the denial report draw the line for you.
A front-desk workflow automation playbook for spas and beauty: which tasks to automate first with AI voice and chat agents to cut admin and capture revenue.
Automate patient intake and appointment scheduling with AI agents that collect details, book in your EHR, and confirm visits. A 2026 guide for clinics.
Automate dental insurance verification and patient intake with AI — collect details, check eligibility, and write clean records to your PMS before each visit.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI