By Sagar Shankaran, Founder of CallSphere
Split the prior-auth queue: fast model for clean continuations, strong model for biologic starts and appeals - with the 91-hours-a-month arithmetic behind it.
Key takeaways
Most rheumatology practices that bought an automated prior-authorization tool two years ago tell the same story. It filled the request competently for a bone density scan. Then it hit a first-time biologic start, misread a step-therapy history spanning two prior carriers, asserted the patient had failed methotrexate when the chart showed intolerance at eight weeks, and the prior-authorization specialist had to unpick the whole thing before submitting. Within a month she had stopped using it for anything that mattered — which meant using it only on the requests that were never the problem.
That failure was not about bad technology. It was about running one setting across two completely different jobs. The 2026 standard is boring in the best possible way: routine requests go to a fast, cheap model, and only genuinely hard ones escalate to the strong one. Cisco built exactly that into the personal AI agent it is rolling out to roughly 90,000 employees, balancing cost against capability request by request. A five-physician rheumatology group can draw the same line in an afternoon.
Model routing means the practice decides in advance which requests get the fast cheap model and which ones get the expensive careful one, based on a written rule about the work — not on who happens to be at the keyboard.
Every practice administrator knows this in their bones but rarely writes it down. The prior-auth work queue contains two different animals.
The routine queue. A follow-up bone density scan on an established osteoporosis patient. A joint ultrasound. A methotrexate or hydroxychloroquine refill authorization. A continuation of an infusion the patient has tolerated for three years, where the chart evidence is clean and the payer's criteria are published, so the answer is a matching exercise between what the policy asks for and what the chart already says. Roughly three-quarters of the volume and almost none of the anguish.
The hard queue. A first biologic start where the step-therapy history has to be assembled from records that predate your chart system and the patient's current plan. A denial with a medical-necessity reason code and a fourteen-day appeal window. A switch between agents after a serious infection. Off-label use. Anything that is not reporting a fact from the chart but making an argument in the physician's clinical voice. Maybe a quarter of the volume and essentially all of the revenue risk, because a biologic stalled five weeks in authorization is a patient flaring, an infusion chair sitting empty, and a drug you bought and cannot bill for.
Timing tightened this year too. The federal prior-authorization rule effective in January shortened decision windows at Medicare Advantage and Medicaid plans to 72 hours for expedited requests and seven calendar days for standard ones. Faster payer decisions only help if your side of the paperwork is ready — and January is exactly when every biologic patient's authorization resets with the new plan year.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for healthcare in your browser — 60 seconds, no signup.
The rule that works in this specialty is not about the drug and not about the dollar amount. It is about one question: does this request assert medical necessity in the physician's words, or does it report facts already in the chart?
Reporting facts is a matching job. The payer's published policy wants a diagnosis code, a documented trial of a preferred agent, a lab value inside a window, and a date. That is well within a fast cheap model, and it is three-quarters of the queue. Making an argument is different: it needs the whole chart read at once, it has to cite where each claim came from, and it will end up in front of a payer's medical director. That one is worth the expensive model every time, and it still gets a physician's signature.
flowchart TD
A["Auth request lands in the work queue"] --> B{"Does it argue medical necessity in the physician's words?"}
B -->|No| C["Fast model drafts from the payer policy and the chart"]
B -->|Yes| D["Strong model builds the packet with chart citations"]
C --> E{"Did the payer policy criteria match cleanly?"}
E -->|Yes| F["Prior-auth specialist reviews and submits"]
E -->|No| D
D --> G["Physician reads, edits and signs, then specialist submits"]
F --> H["Outcome and denial code logged"]
G --> H
Note the arrow from the routine lane back into the hard lane. That is the part the 2024 tools got wrong. A routine request that does not match the policy cleanly is no longer routine — it has just told you it is hard — and it should change lanes automatically instead of landing on the specialist's desk half-finished.
The routing rule should be written by the prior-authorization specialist and the billing manager together, on a single page of plain English, and signed off by the physician-owner. The specialist knows which payers publish usable criteria and which hide them. The billing manager knows which denial codes are costing money this quarter. The physician-owner has one job in that meeting: naming the categories that may never be drafted by the cheap model, whatever the volume looks like.
Then review it monthly against the denial log, not against a feeling. If routine-lane requests come back denied at a higher rate than the hard lane, the line is in the wrong place — usually because a payer quietly updated its policy. That review takes twenty minutes and belongs in the meeting where the practice already goes through aging accounts receivable.
One warning: never let the person drowning in the queue decide what counts as routine. At 4:45 p.m., everything looks routine. That is why the rule is written in advance, by two people, and approved by a third.
Assumptions, stated and illustrative: five rheumatologists, an in-office infusion suite, roughly 620 authorization and re-authorization requests a month, split about 74% routine and 26% hard. Prior-authorization specialist time is costed at a fully loaded $27 an hour.
| Queue | Requests / month | Minutes each, by hand | Minutes each, with routing | Hours / month after |
|---|---|---|---|---|
| Routine (scans, refills, clean continuations) | 459 | 19 | 6 | 45.9 |
| Hard (first biologic start, appeals, switches) | 161 | 19 | 22 | 59.0 |
| Total | 620 | 196 hours | — | 104.9 hours |
That is 91 hours a month back, roughly $2,457 of specialist time at the assumed rate. The model cost of getting there is the small number: at illustrative prices, 459 routine drafts on the fast model plus 161 packets on the strong one lands near $200 a month. Run everything on the strong model instead and it is closer to $590 — still not the expensive part, which is exactly why owners get the trade backwards. Routing is not primarily about saving money on the AI bill. It is about the hard cases getting the careful treatment while the routine ones stop consuming a specialist's whole morning.
The real return sits elsewhere on the profit and loss statement: ninety-one hours is most of a part-time position you do not have to add in January, when plan-year turnover hits and every established infusion patient needs a fresh authorization inside three weeks.
Still reading? Stop comparing — try CallSphere live.
See the healthcare AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
The peer-to-peer call. When a payer's medical director wants to talk, that is a physician-to-physician conversation about a specific patient, and no drafting changes who has to be on the phone. What drafting can do is put the chart evidence, the failed agents with dates and reasons, and the payer's own policy language on one page before the call — often the difference between winning it and rescheduling it.
Anything that touches step-therapy history from outside records. If the evidence that the patient failed two preferred agents lives in a fax from a rheumatologist in another state, a person needs to read that fax. Getting this wrong is not a paperwork error; it is a false statement to a payer with the practice's name on it.
And any appeal with a hard deadline gets a human owner with a name and a date, not a queue. Fourteen-day and thirty-day appeal windows do not forgive a workflow that assumed somebody would notice.
Most practices set it up inside the tool they already have: a default model for the staff, plus the ability for the prior-auth specialist to escalate a single request. If your vendor cannot tell you which model handled a given request, ask that before renewal.
Watch the denial log by lane for three months. Routine-lane requests should be denied at or below the rate your specialist achieved by hand. If it climbs, either the line moved or a payer changed its policy — both of which you want to catch in the log rather than in a patient's flare.
Dollar value is a poor routing test on its own, but a fine reason to add a review step. Many groups keep high-cost continuations in the routine lane for drafting and require a second read by the specialist before submission. Cheap to draft, carefully checked, submitted by a person.
It is the single place it helps most, because January is almost entirely routine volume arriving at once — established patients, unchanged therapy, new plan card. Build the rule in October so it is running before the plan year turns, not in the second week of January with the queue already three hundred deep.
Pull last month's authorization log and mark every line with one letter: R if the request reported facts from the chart, H if it made a clinical argument. Do not overthink the edge cases; you want the shape of the split, not perfection. That sorted list, taken to the billing manager and the physician-owner, is the routing rule in draft form, and it takes the specialist about an hour.
The authorization queue also generates its own phone traffic — patients asking whether the approval came through, whether Thursday's infusion is still on, whether the specialty pharmacy shipped. Those calls land on the same front desk already answering new-referral calls. CallSphere builds AI voice and chat agents that answer the phone and web chat, book appointments and capture new patient enquiries around the clock, so the person handling authorizations can stay on authorizations.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Sixty-one pages of addendum land three days before a DOT letting. How model routing gets an estimator the quantity changes that actually move the bid.
Model routing sends routine product requests to a cheap model and hard ones to a strong one. Here is who draws the line in a B2B software company, and how.
Cheap model for potholes, strong model for water quality, a pager for sewage in a basement. How a public works superintendent writes the triage table.
Prior-authorization phone tag drains Madison psychiatry practices. An AI agent absorbs payer callbacks, pharmacy rejections, and patient status-check calls.
How water systems use a cheap-model first pass and a strong-model escalation so nobody drives out at 2 a.m. for a SCADA alarm that already cleared itself.
Sort courier calls before answering: cheap model for status and reschedules, strong model for damage claims and missed medical runs. Costs, rules and limits.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI