By Sagar Shankaran, Founder of CallSphere
How a therapy practice builds a 180-call test set from its own history, what a passing score looks like on risk disclosures, and how to widen authority in stages.
Key takeaways
You already tried this. Sometime in 2024 a vendor sold your practice a phone bot, it answered "what are your hours" correctly and then told a woman whose son had stopped eating that the next opening was in eleven weeks, and your clinical director killed it inside a month. Fair. Here is what is actually different in 2026, and it is not that the agents got smarter — although they did.
What changed is that you can now watch one work, step by step, on calls you already have the answers to, and refuse to widen its authority until the score is good enough. The tooling that measures agents against real cases and shows what the agent did at each step matured in 2026 into the thing that decides whether a rollout happens at all. In behavioral health that matters more than in almost any other trade, because the failure mode is not a lost sale.
Every outpatient practice has an intake funnel that leaks. A referral arrives — from a PCP's office, a school counselor, a hospital discharge planner, a Psychology Today profile, a parent Googling at 11 p.m. — and hits either your intake coordinator or your voicemail. She calls back, twice, gets voicemail, marks it in the spreadsheet, and by day four the person has called somebody else. Practices that have measured it usually find a third of inbound referrals never convert to a first appointment, and most of those die in the callback gap rather than in the clinical fit conversation.
So handing intake to an agent is obviously attractive. It is also the single riskiest thing you could hand it, because somewhere in that call volume is a caller who says "I don't think I can keep myself safe tonight." An agent that answers the insurance question beautifully and misses that sentence is not a partial success. It is the thing that ends your practice.
Before an intake agent talks to a single new caller, grade it against a set of real past calls from your own practice where you already know what the right handling was — and only widen what it is allowed to do when the score on the dangerous cases is perfect. That is the whole discipline, and every practice can build the set, because you already have the raw material.
Pull 180 intake contacts from the last year: recorded calls if you have them, voicemail transcripts, the intake coordinator's notes, the web form submissions. Then sit with your clinical director and label each one with what should have happened. You want a deliberately unbalanced set:
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for behavioral health in your browser — 60 seconds, no signup.
flowchart TD
A["180 labelled past intake calls"] --> B["Agent answers each one in test mode"]
B --> C{"Did it catch every risk disclosure?"}
C -->|No| D["Clinical director reviews the transcript step by step"]
D --> E["Tighten the escalation rule"]
E --> B
C -->|Yes| F["Live on scheduling and insurance questions only"]
F --> G["Weekly score on real calls"]
G --> H["Widen authority one task at a time"]
Two scores, held to completely different standards. On risk disclosure, the target is every single one, caught and warm-transferred to a human or routed to your crisis protocol — 15 of 15, not 14. A miss on that set means the agent does not go live on intake at all, no matter how well it did on the other 165. And you also count the opposite error: how often it escalated something that did not need it. A few false alarms are fine; your on-call clinician can absorb them. Fifty a week is a different problem, because that is how a crisis protocol gets ignored.
On everything else — insurance, availability, scheduling — you want accuracy in the high nineties, and specifically you want to know what it says when it does not know. "Let me have someone call you back today about that" is a correct answer. Inventing a copay is not, and inventing an availability date is worse, because the client shows up.
Illustration only. Assume 320 inbound intake contacts a month. Assume your intake coordinator currently converts 62% of them to a booked first appointment, and that from your own recorded history roughly 2% of contacts contain a risk disclosure — about six a month.
| Measure | Today | Agent at the score you should demand | Agent at a score you should reject |
|---|---|---|---|
| Risk disclosures caught, monthly | 6 of 6 | 6 of 6 | 5 of 6 |
| Contacts answered live, first attempt | 71% | 100% | 100% |
| Booked first appointments | 198 | ~230 | ~230 |
| Added sessions in month one at $118 allowed | — | ~$3,800 | ~$3,800 |
| Acceptable? | — | Yes | No, at any revenue number |
The last row is the argument. The revenue from filled slots is identical in both agent columns, which is exactly why the revenue number must not be the one you decide on. The 180-call test set exists to let you tell those two columns apart before a real caller does.
Nobody should go from a graded test to full intake in one step. The sequence that works: first, the agent answers hours, location, insurance panels and self-pay rates, and takes a message for everything else — two weeks, with your intake coordinator reading every transcript on Friday. Second, it books into open slots for established clients doing a reschedule, which is low-stakes and high-volume. Third, it books new-client intakes for the clinicians whose caseloads you actually want filled. Fourth, and only if the numbers hold, it handles the after-hours line where today the practice sends everyone to voicemail with a recording pointing to 988.
At every stage, keep reading transcripts. Not a sample — all of them, for the first fortnight of each stage. It is an hour a week for your practice manager and it is the entire safety case.
The agent does not perform clinical triage, does not assign a risk level, does not make a hospitalisation decision, and does not do safety planning. Its only job on a risk call is to recognise it inside the first few sentences, stay warm and human about it, and get a licensed person on the line or execute your written crisis protocol. Write that protocol first — who is on call, what the transfer path is at 2 a.m., what happens if nobody picks up — because an agent cannot follow a protocol you have not written.
Still reading? Stop comparing — try CallSphere live.
See the behavioral health AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
It also does not fix a real capacity problem. If your next intake slot is genuinely nine weeks out, answering the phone faster mostly means telling people that faster. That is still better than voicemail, and an honest waitlist plus a referral to two practices that do have room protects your reputation with the PCP office that sent them. But do not buy an agent expecting it to solve a staffing shortage.
Finally, be careful with children. A parent calling about a 14-year-old raises consent questions, mandated reporting questions, and in some states questions about the minor's own right to consent to treatment. Keep those calls on the human path until you have specifically tested that path and your clinical director has signed off on it.
Yes. Use voicemail transcripts, the intake coordinator's call notes, and web form submissions, and have her write out from memory the ten calls that scared her most in the last year. A test set of 60 well-labelled cases beats 400 unlabelled ones. Check your state's recording consent rules before you start recording new calls.
Your clinical director, with the intake coordinator in the room, and it takes about four hours for 180 cases. Do not delegate the risk labels to an administrator and do not let a vendor label them for you — the labels are the standard you are holding the agent to, and they need to be yours.
Monthly, and always after the vendor changes their model. That last part matters more than owners expect: the underlying models get updated regularly now, and an agent that scored 15 of 15 in May is not automatically scoring 15 of 15 in July. Keep the test set, re-run it, keep the scores in a file your board or your malpractice carrier could read.
Ask them before you go live, not after. Carriers in this space are actively working out their position on AI-assisted intake, and the practices that have an easy conversation are the ones who can produce a written crisis protocol, a graded test set, and a log showing every risk call was transferred to a human.
CallSphere builds AI voice and chat agents that answer practice phone lines and web chat, book appointments, and capture referral details 24/7 — and the reason this post is mostly about testing rather than features is that in behavioral health, testing is the product decision. Ask any vendor, including us, to run against your own labelled calls before you point a phone number at them. If they will not, that answers the question.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Sixty calls a day, one person at the counter. See what unanswered banner and yard-sign calls cost a print shop, and what an instant-answer voice agent changes.
Traverse City therapy practices face winter intake waves and waitlist calls no clinician can answer mid-session. See how AI answering captures every inquiry.
Bend counseling practices are full — and still losing clients to missed calls. How an AI agent keeps a live waitlist and refills openings in days, not weeks.
Fayetteville therapists ride the University of Arkansas semester clock. How an AI intake agent answers the student, parent, and referral calls it brings.
Solo therapists in Santa Fe wear every hat — receptionist, scheduler, biller. How an AI voice and chat agent takes over the phone jobs and returns your breaks.
Duluth winters cancel therapy sessions by phone at 7 a.m. How AI answering turns storm-morning cancellations into same-call rebookings and telehealth swaps.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI