---
title: "Score the AI Against 150 Real Intake Calls From Your Own Schedule Before It Books a Single PT Eval"
description: "Before an AI answers your PT clinic's intake line, build a 150-call answer key from last year's schedule. What to sample, how to score, where humans stay."
canonical: https://callsphere.ai/blog/score-the-ai-against-150-real-intake-calls-from-your-own-schedule-befo
category: "Healthcare"
tags: ["physical therapy", "ai agent evaluation", "patient intake", "outpatient rehab", "front desk automation", "webpt"]
author: "CallSphere Team"
published: 2026-06-17T15:56:44.000Z
updated: 2026-07-25T23:16:06.817Z
---

# Score the AI Against 150 Real Intake Calls From Your Own Schedule Before It Books a Single PT Eval

> Before an AI answers your PT clinic's intake line, build a 150-call answer key from last year's schedule. What to sample, how to score, where humans stay.

## 7:52 on a Monday, three lines ringing, and one of them is a post-op ACL

It is 7:52 on a Monday in October at a four-therapist clinic and three lines are lit. Line one is a Medicare patient moving her Thursday to Friday. Line two is a self-pay runner asking what an evaluation costs and whether she needs a doctor first. Line three is the orthopedic office: a post-op ACL reconstruction is being discharged today, needs to start Wednesday, and the script will be faxed "sometime this afternoon." The patient access coordinator has one headset, a filling waiting room, and a therapist at the desk asking whether the 8:00 eval showed.

Every AI voice vendor points at this hour, and they are right about the hour. They are usually wrong about what happens next. Whether an intake agent helps your clinic or quietly bleeds cases out the back has little to do with how natural it sounds. It has everything to do with whether it knows that the ACL on line three cannot be treated under Medicare Part B without a certifying physician signature on the plan of care inside 30 days, and that the runner on line two can be seen under your state's direct access rule while a Medicare beneficiary cannot. None of that shows up in a demo. All of it shows up in your denial report ninety days later.

## What actually changed in 2026: you can grade the agent before it speaks to a patient

Through 2024 and 2025, buying an AI phone agent meant sitting through a demo, liking it, switching it on, and finding out. The tooling that grew up in 2026 changed the order of those steps. Evaluation and observability — scoring an agent against a fixed set of real past cases, then watching each live conversation step by step to see what it looked up and what it decided before it spoke — became the thing that gates a rollout rather than a nicety bolted on afterward. You measure, you watch, and only then do you widen what the agent may do without a person in the loop.

**An evaluation set — in plain English, an answer key — is a fixed list of real past cases from your own clinic where you already know the correct outcome, used to grade an AI agent before it is ever allowed to speak to a patient.**

The good news for a rehab practice is that you are already sitting on the raw material. Your history in WebPT, Prompt, Raintree, Jane, TheraOffice or Net Health records what actually happened to every intake call last year: the payer, whether an authorization was required and obtained, how many visits were approved, whether a referral was on file, which clinician the patient landed with, and whether the case survived past visit three. Nobody has to invent test cases. You have twelve months of them, with the right answers already written down by the people who did the work.

## Building the 150-call answer key out of last year's schedule

Pull a stratified sample, not a random one. A random 150 calls from a busy clinic is ninety percent easy reschedules and teaches you nothing. You want the mix that actually breaks front desks:

- **30 clean commercial PPO evaluations** — referral on file, benefits verified, no authorization. The agent must get these perfect; they are the volume.
- **25 Medicare Part B new patients** — the certification and plan-of-care signature question, plus anyone near the annual therapy threshold where the KX modifier starts to matter.
- **25 Medicare Advantage cases with a delegated vendor** — plans that hand utilization management to a third party and approve visits in blocks of six.
- **20 workers' compensation calls** — an adjuster, a claim number, a nurse case manager wanting the eval, a state fee schedule.
- **15 auto and no-fault calls** — personal injury protection, an attorney's office in the middle, a letter of protection.
- **20 cancellation and no-show recovery calls** — where filling the slot off the waitlist is the whole job.
- **15 genuinely ugly calls** — acute calf swelling two weeks post-op, a parent booking a pediatric eval, an angry caller about a balance after the deductible reset.

For each, the office manager and the billing specialist write the correct outcome in a spreadsheet: what the agent should have said, what it should have booked, what it should have refused to answer. That spreadsheet is the deliverable. Two people, about six hours across a week, and it outlives any vendor you sign with.

```mermaid
flowchart TD
  A["Pull 150 real intake calls from last year's schedule"] --> B["Office manager and biller write the correct answer for each"]
  B --> C["Agent runs all 150 with the phone line still off"]
  C --> D{"Payer, referral rule and eval slot all correct?"}
  D -->|Below 95 percent| E["Read the step-by-step record, fix the rules, rerun"]
  E --> C
  D -->|95 percent or better| F["Answer overflow only, 7 to 9 a.m."]
  F --> G["Clinic director reviews 20 live calls every Friday"]
  G --> D
```

## Scoring: which mistakes are free and which cost you a case

Do not score the agent on a single percentage. A rehab intake call has four separable decisions with wildly different consequences. Grade each separately.

**Identification** — did it get the payer, the body part and the referring provider right? A miss here is usually recoverable at the desk. **Eligibility** — did it correctly decide whether a referral or a certification was needed before the first visit? A miss here produces a visit you cannot bill. **Booking** — did it put the patient with a clinician who can legally and clinically treat them, at a time your schedule supports? A miss here burns an evaluation slot, the scarcest thing most clinics own. **Restraint** — did it decline to answer things it should not? An agent that confidently quotes a remaining deductible from stale eligibility data costs more in goodwill than it ever saved in labor.

The step-by-step record is what makes this fixable rather than mysterious. When the agent tells a Medicare Advantage patient she is approved for twelve visits, you can open that conversation and see which eligibility check it ran and where it drew the wrong conclusion. In 2024 you got a transcript and a shrug.

## The arithmetic on a bad booking

Why the answer key pays for itself, in round numbers you should replace with your own. All figures below are illustrative assumptions, not measured results.

| **Assumption** | **Value** |
| --- | --- |
| New evaluations booked per month | 60 |
| Average visits per completed case | 9 |
| Average net collection per visit | $95 |
| Revenue per completed case | $855 |
| Intake calls the agent handles unsupervised | 60% |
| Bad-booking rate before evaluation and tuning | 4% |
| Bad-booking rate after tuning against 150 cases | 1% |
| Share of bad bookings that never rebook | 50% |

Sixty evals at sixty percent is 36 calls a month the agent owns. At four percent error that is 1.44 bad bookings; at one percent, 0.36. The difference is roughly one a month, half of which walk permanently: about 0.54 lost cases, or $462 a month, $5,540 a year — before the denied visits you write off and the therapist hour that sat empty. Against six hours of office-manager time and a few hours a month of call review, the math is not close, and the answer key gets reused every time a payer changes its rules.

## Where the answer key runs out and a human has to pick up

An evaluation set only covers what already happened. It cannot grade the agent on a payer policy that changes next Tuesday, or a patient describing symptoms that belong in an emergency department rather than on your Wednesday schedule. Three things stay with people no matter how good your scores get.

First, **anything that sounds like a red flag**. Unexplained calf swelling after a knee replacement, new bowel or bladder changes with low back pain, chest pressure on exertion — the agent's only correct move is to stop, say it cannot advise, and transfer to a licensed clinician or tell the caller to seek immediate care. Score restraint hard on those fifteen ugly calls.

Second, **anything that commits money**. Quoting a remaining deductible in January, when half your patients have reset and your eligibility data is a week stale, is how a clinic ends up eating balances. Third, **the discharge and plan-of-care conversation**. Recertification at 90 days, the progress note at the tenth visit, whether a plateaued patient should be discharged — clinical judgments belonging to the treating physical therapist, as your state practice act says plainly.

## Frequently asked questions

### How many calls do I actually need in the answer key? Is 150 a real number or a round one?

A working number, not a rule. What matters is that every category has 15 to 20 examples; below that, one wrong answer swings the score more than the truth warrants. A single-site clinic with two payers gets away with 80. A three-site practice taking comp, no-fault and two Medicare Advantage plans wants closer to 250.

### Can I use recorded patient calls for this without a HIPAA problem?

You can use your own recordings and scheduling history for your own operations, but a vendor doing the scoring is handling protected health information, so you need a signed business associate agreement before anything leaves your building. Many practices sidestep it by de-identifying the answer key — case numbers instead of names and dates of birth — since the agent is graded on the decision, not the identity.

### What score is good enough to let it answer the phone unsupervised?

Set the bar by consequence, not by an average. Most clinics land near 98 percent on eligibility and restraint, 95 percent on booking, and whatever you get on identification, since the desk catches those. If eligibility sits at 90 percent you are not ready, no matter how the rest scores.

## The one thing to start on Monday

Do not shop for a vendor first. Open your scheduling system, export every new-patient intake from the same month last year, and sort it by payer. Within an hour you will know whether your hard cases are workers' comp, Medicare Advantage, or the physician offices that fax scripts late. That distribution is the first column of your answer key.

[CallSphere](https://callsphere.ai) builds AI voice and chat agents that answer clinic phone lines and web chat, book appointments and capture new-patient inquiries around the clock — including the 7:00 a.m. overflow and the Saturday calls that currently reach voicemail. If you are evaluating any agent for your intake line, ours included, build the answer key first and make the vendor score against it. A vendor who will not be measured against your own history is telling you something.

---

Source: https://callsphere.ai/blog/score-the-ai-against-150-real-intake-calls-from-your-own-schedule-befo
