By Sagar Shankaran, Founder of CallSphere
Why a payer records request eats five weeks of one person's time, and what changes for the calendar when several AI agents review all 200 charts at the same time.
Key takeaways
The envelope is thin, which is always a bad sign. It is from the managed care organization that administers your state's Medicaid behavioral health benefit — Carelon, Optum, Magellan, whichever one holds your largest contract — and it names 200 dates of service across 14 months. Records due in 30 days. Your compliance manager, who is also your utilization review person and also the one who chases missing supervisor signatures, does the arithmetic out loud in the doorway: she can review about eight charts a day properly. Forty a week. Five weeks. She has four.
So the practice does what every practice does. She spot-checks. She pulls the charts for the clinicians she worries about, skims the rest, and ships the box. Then everyone waits four months to find out which notes the auditor decided did not support the code.
Not quality of care. Documentation. Specifically the thread that has to run unbroken from the intake assessment through the treatment plan to every progress note: the diagnosis on the claim matches the diagnosis in the assessment, the treatment plan names a goal and an objective that the session actually addressed, the note says what the clinician did rather than what the client talked about, the start and stop times support the code billed, the plan was reviewed before it expired, and the right people signed in the right order — the associate-level clinician and the supervising LCSW or licensed psychologist, both, within the window your state board allows.
The failures are boring and expensive. A 90837 billed on a session documented as 48 minutes. A treatment plan that expired in March with notes running to July. An LPC-Associate's note with no countersignature. A crisis contact documented in a phone log but never in the chart. Each one is a recoupment of the full session, and if the sample rate is bad enough the payer extrapolates across the whole period.
Nothing about reading one chart is hard. The problem is that 200 charts is one queue with one person at the front of it, and every chart requires opening several documents in SimplePractice or TherapyNotes or Valant, cross-checking the claim in the billing module, and holding six rules in your head at once. It is serial work performed by a role you have exactly one of. That is why it eats five weeks: not difficulty, order.
Agent Teams, released as a research preview alongside Claude Opus 4.6, lets several AI agents split one large job between them and work at the same time, then merge what they found into a single result — which turns a queue of 200 charts into a job whose length depends on the pile, not on the person. That is the entire change, and for a behavioral health practice it lands squarely on audit prep, treatment-plan expiry sweeps, and the pre-billing note review nobody has time to do.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for behavioral health in your browser — 60 seconds, no signup.
flowchart TD
A["Audit letter: 200 dates of service"] --> B["Pull charts and claims from the EHR"]
B --> C1["Agent 1: charts 1-50"]
B --> C2["Agent 2: charts 51-100"]
B --> C3["Agent 3: charts 101-150"]
B --> C4["Agent 4: charts 151-200"]
C1 --> D["Merged exception list by clinician and rule"]
C2 --> D
C3 --> D
C4 --> D
D --> E["Compliance manager works the 31 flagged charts"]
Your practice manager exports the date-of-service list from the billing module and hands the whole job over as one instruction: for each of these 200 encounters, check the billed code against the documented duration, confirm the diagnosis matches the current assessment, confirm an active treatment plan covered the date, confirm the note references a goal from that plan, confirm the signature and countersignature exist with dates, and flag anything missing with the exact chart, date, clinician and rule.
Four agents take 50 charts each. They finish before lunch. What comes back is not a pile of documents — it is an exception list: 31 charts with a problem, sorted by clinician, with the rule that failed named in plain language. Nine are countersignatures missing on one associate's caseload from a two-week stretch in February when her supervisor was on leave. Six are treatment plans that lapsed. Four are duration-versus-code mismatches on the same clinician, which is a training conversation, not a fraud conversation. The rest are one-offs.
Your compliance manager now spends her four weeks doing the part only she can do: deciding what is fixable through a late entry that is properly dated as such, what has to be voided and rebilled at the lower code, what gets self-disclosed, and which clinician needs an hour of supervision on documentation before September's referral surge.
Illustration only — use your own numbers. Assume a group practice with 14 clinicians, an average allowed amount of $118 per session, and a 200-encounter audit sample. Assume the manual spot-check catches half of the documentation defects before submission and the parallel review catches most of the rest.
| Line | Spot-check today | Parallel review |
|---|---|---|
| Charts actually reviewed before submission | 60 of 200 | 200 of 200 |
| Compliance manager hours consumed | ~48 | ~14 |
| Defective encounters shipped to the payer | 18 | 6 |
| Recoupment at $118 each | $2,124 | $708 |
| Extrapolation exposure if the error rate crosses the payer's threshold | real | unlikely |
The recoupment line is not the point. Thirty-four hours of your compliance manager's time is roughly $1,400 loaded, and the extrapolation risk is the number that actually keeps owners awake — a 9% error rate applied to 14 months of a Medicaid panel is a five-figure event, and it arrives as a demand letter, not a negotiation.
Once the audit is answered, the useful move is to point the same parallel review at the whole active caseload once a month: every treatment plan expiring in the next 45 days, every note older than the countersignature window, every client whose PHQ-9 has not been re-administered in a quarter, every authorization approaching its last unit. That sweep has always been theoretically possible and practically impossible, because it was one person reading charts one at a time. Now it runs overnight on the first Sunday of the month and lands in the practice manager's inbox as a to-do list with names on it.
An agent can tell you a note does not reference a treatment plan goal. It cannot tell you whether the clinical decision was right, whether the client's risk level justified the frequency, or whether a late entry is honest housekeeping or something you should be reporting. Do not let it write clinical content into a chart. A note it drafts is a draft the clinician reads, edits and signs — the signature means the clinician stands behind it, and that does not change.
Still reading? Stop comparing — try CallSphere live.
See the behavioral health AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
Two more limits worth naming. First, if your practice holds substance use disorder records covered by 42 CFR Part 2, the consent rules for disclosing those records are stricter than HIPAA and your business associate agreement with any AI vendor needs to say so explicitly before a single chart moves. Second, the exception list is only as good as the rules you hand it; if your state's Medicaid manual requires a specific treatment plan review interval and you did not say so, the review will not catch it. Write the rule list with your compliance manager, from your actual payer contracts, and version it.
Do not start with the audit. Start with 25 charts you already know the answers to — ideally ones from a past audit where the payer told you exactly what was wrong. Run the review on those 25 and compare it to the findings you already have. If it catches what the auditor caught, you have a measured tool. If it misses two, you learn which rules you failed to write down. That takes an afternoon and it is the only honest way to know whether the thing works on your charts, in your EHR, with your clinicians' writing habits.
Not necessarily, and this is the question to put to any vendor in writing. Ask where records are processed, whether they are retained after the job, whether a business associate agreement covers it, and whether an on-premises option exists. Practices with Part 2 records or county contracts that specify data handling should assume the answer must be yes to the last one.
It is the real obstacle, more than the AI. SimplePractice, TherapyNotes and Valant all have reporting and export paths, but the claim data and the note data usually live in different reports, and matching them is fiddly the first time. Budget a day of your billing specialist's time to build the two exports once. After that it is a repeatable job.
No, and if a vendor tells you otherwise, they have never sat through a payer audit. It replaces the reading, not the judgment. Practices that try this generally find their compliance person becomes more valuable, because she is finally working on the 31 charts that matter instead of skimming 200.
Same shape of problem, and it is arguably the better first target: a parallel sweep of unsigned notes older than 72 hours, sorted by clinician, sent as one message to each clinician's supervisor on Friday morning. It is lower stakes than an audit and it tells you within two weeks whether your team will actually act on the list.
Audit season and September's referral surge tend to arrive within a few weeks of each other, and the first casualty is the intake line — calls that go to voicemail while the front desk is pulling records. CallSphere builds AI voice and chat agents that answer the practice phone and web chat around the clock, take intake details, and book the first available appointment, so a new referral does not sit in a voicemail box while your team is answering a payer's records request. It does not touch your charts or your audit response; it keeps the front of the practice open while the back of it is buried.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
A 40-claim DME probe takes 16 working days one at a time and finds the gaps too late. Split four ways, the records requests go out on day two of forty-five.
Nobody built a connection between veterinary software and the state monitoring portal. Computer use closes that gap - with the limits an owner should insist on.
A 640-line BOM scrub eats three days of your buyer's week. Here is what splitting the RFQ across several agents does to quote throughput at an EMS shop.
Map one messy process and prove it: spray ticket records, the five baseline numbers to capture before you start, and the error rate that ends the debate.
Reno's population is growing faster than its behavioral-health capacity. How psychiatry practices use an AI voice and chat agent to answer every call, 24/7.
Spokane behavioral-health groups serve the whole Inland Northwest. How an AI voice agent stops referral leakage, status-call floods, and clinic phone tag.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI