By Sagar Shankaran, Founder of CallSphere
Sort courier calls before answering: cheap model for status and reschedules, strong model for damage claims and missed medical runs. Costs, rules and limits.
Key takeaways
Most couriers who tried this the first time have the same story. You bolted something onto the delivery-status line, it answered the easy ones badly and the hard ones catastrophically, and within a month your dispatcher had trained every regular consignee to mash zero. The customer service rep ended up handling the same volume plus the cleanup from what the bot told people. You turned it off and told the vendor's rep not to call back.
What changed is not that the models got smarter in a way that fixes the hard calls. It is that in 2026 the sorting happens first. A cheap, fast model takes the routine contact, and only the genuinely hard one is handed up to an expensive model or a person. Cisco built exactly this into the personal AI agent it is rolling out to roughly 90,000 employees — route by difficulty, hold the cost down, keep the capability available for when it is actually needed. The same logic scales down to a twenty-two-van operation with two people on the phones.
Model routing means every incoming contact is sorted before it is answered: the fast cheap model handles the routine ones, and only the cases that meet rules you wrote get escalated to a stronger model or a human.
Look at last week's inbound. In a last-mile operation the first pile is enormous and boring: where is my package, was it delivered, can I get the photo again, the driver could not find the gate code, can we move it to Thursday, my apartment office closes at five. These are answerable from the scan history and the stop record. There is exactly one right answer and it is already in your system.
The second pile is small and expensive. A $4,200 range arrived with a crease in the door and the consignee wants a claim opened. A 2 a.m. STAT run to a hospital lab did not happen and the lab manager is on the line. A driver clipped a mailbox. A shipper's account manager is asking why the completion rate on their route slipped two days running before a contract review. Those calls have money, a scorecard, or a relationship attached, and a wrong sentence in the first thirty seconds costs you far more than the call.
Every courier already routes these — badly, by whoever picks up. The dispatcher who is best at claims is also the one releasing routes at 6:40 a.m., so at 6:40 a.m. the claim gets whoever is free.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for logistics in your browser — 60 seconds, no signup.
The routine pile is now handled end to end by a fast speech-to-speech model — something in the Gemini 3.1 Flash Live class, which answers in roughly the time it takes a person to draw breath and can look up a stop mid-sentence rather than making the caller wait. It reads the scan history, gives the delivery window, texts the proof-of-delivery photo, moves the stop, writes the gate code into the stop notes so the driver sees it on the next attempt.
flowchart LR
A["Call or web chat hits the dispatch line"] --> B{"What kind of contact is this?"}
B -->|"Where is my package"| C["Fast model reads scan history and answers"]
B -->|"Reschedule or gate code"| D["Fast model updates the stop on the route board"]
B -->|"Damage, theft, missed STAT run"| E["Strong model drafts the claim and pulls POD photos"]
E --> F["On-road supervisor approves before anything is promised"]
C --> G["Logged to the account's daily exception report"]
D --> G
F --> G
The second pile gets handed up. A stronger model — Claude Sonnet 5, GPT-5.6, whatever your shop settles on — assembles the case before a human touches it: the scan trail, the driver's photo at the door, the geostamp, the account's damage-claim terms, the shipper's notification deadline. Your on-road supervisor picks up a call that is already documented instead of one that starts from nothing. Nothing is promised to the caller until she approves it.
This is the part vendors get wrong and owners have to own. The dividing line is not a technical setting. It is a business rule, and it belongs to you and your operations manager.
In practice a courier's line has four tripwires. Money: any contact mentioning damage, loss, or a dollar figure over your claim threshold escalates immediately. Account tier: your top five shippers by revenue escalate on anything beyond a status question, because a two-minute save on a $600,000 account is not a saving. Cargo type: medical, controlled, temperature-controlled and anything moving under an air-cargo security program never gets a cheap answer, full stop. And repetition: a third contact on the same tracking number in twenty-four hours is by definition not routine, whatever the caller says it is about.
Write those four rules on one page, sign them, and revisit them monthly. Claude's enterprise governance update on 2 July 2026 added spend limits per person and per team with alerts at 75% and 90% of budget, plus a usage dashboard — which is how you find out on the 12th that your escalation rule is firing on 40% of calls instead of the 15% you designed for, rather than finding out on the invoice.
Illustration, not a benchmark. Assume 4,000 inbound contacts a month across phone and web chat, a dispatcher's fully loaded cost of $27 an hour, and 3.5 minutes average handling time.
| Contacts a month | 4,000 |
| Routine share (status, POD resend, reschedule, gate code) | 82% = 3,280 |
| Hard share (damage, missed medical run, top-tier account, repeat contact) | 18% = 720 |
| Human cost of the routine pile today, at 3.5 min and $27/hr | $5,166 a month |
| Model cost, routine pile, fast model | about $66 a month |
| Model cost, hard pile, strong model preparing the case | about $101 a month |
| Human time still spent, hard pile only, at 9 min each | $2,916 a month |
| Difference | roughly $2,080 a month, plus 191 hours off the desk |
Notice which line is small. The model bill is not the story — sending every one of those 4,000 contacts to the most expensive model would still only cost a few hundred dollars a month, because frontier pricing has come down roughly tenfold since 2025. The reason to route is speed on the routine calls and undivided attention on the expensive ones. The 191 hours land mostly between 6:30 and 9:30 a.m., which is when your dispatcher is otherwise trying to release routes with a phone against her ear.
Never let the automatic side of this promise money. Not a refund, not a claim outcome, not a credit on the next invoice. It can open the claim, attach the photos, and say a supervisor will call back within the hour. That single rule prevents most of the damage a mis-routed call can do.
Still reading? Stop comparing — try CallSphere live.
See the logistics AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
Keep a person on anything involving a driver incident. If a caller says a van hit something, someone was rude at the door, or a driver was in a place he should not have been, that is a human conversation and a safety file, and the tone of the first thirty seconds affects whether it becomes a claim.
And watch the sorting itself for the first six weeks. Pull twenty calls a week that the router called routine, listen to them, and count how many should have gone up. If it is more than one in twenty, your rules are too loose, not the model. The line moves; you move it.
They will, mostly, and the ones asking where their package is do not care as long as the answer is right and instant. Say so plainly at the start of the call. Where it goes wrong is when the routine model gets caught pretending on a hard call it should never have taken, which is a routing failure, not a disclosure failure.
The routine pile grows far faster than the hard pile in November and December, which is exactly the shape routing handles well. Budget for it: the fast model absorbs a Cyber Monday Tuesday without you hiring seasonal phone help, while the hard pile stays roughly proportional to stops and still needs your supervisor.
Yes. The answering side reads from and writes to whatever you already run — Onfleet, DispatchTrack, Circuit for Teams, Route4Me. The one thing worth checking before you start is whether stop notes and reschedules can be written back automatically, because if they cannot, every reschedule still becomes a message for a person to retype.
Sample. Twenty calls a week, listened to by your operations manager, scored against the scan history. That habit is worth more than any dashboard, and it takes about forty minutes.
Take proof-of-delivery photo resends and delivery-window questions only. Nothing else. Route those to a fast model, leave everything else ringing to the desk as it does today, and compare two weeks of handling time against the two weeks before. That is a small enough change that if it fails, you have lost nothing but the setup afternoon.
CallSphere builds the answering layer for exactly this shape of problem: voice and chat agents that pick up the dispatch line around the clock, handle the routine status and reschedule contacts, and hand the damage claim or the missed medical run to your supervisor with the details already gathered.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Sixty-one pages of addendum land three days before a DOT letting. How model routing gets an estimator the quantity changes that actually move the bid.
Model routing sends routine product requests to a cheap model and hard ones to a strong one. Here is who draws the line in a B2B software company, and how.
Cheap model for potholes, strong model for water quality, a pager for sewage in a basement. How a public works superintendent writes the triage table.
How water systems use a cheap-model first pass and a strong-model escalation so nobody drives out at 2 a.m. for a SCADA alarm that already cleared itself.
Six tests that make a title order routine, the escalation list that never bends, who owns the rule, and a costed 90-file month showing where the money sits.
Model routing for a management office: which tenant emails a fast model can answer, which need a lease read, and the escalation triggers that protect you.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI