By Sagar Shankaran, Founder of CallSphere
Sort six months of front-counter recordings into eight call types, write the right answer for each, then widen the agent's authority in four stages, not one.
Key takeaways
You already tried this. Somewhere around 2024 a vendor sold you a phone system with a menu tree and a bit of voice recognition on the front, and within a month you had a customer telling an adjuster that nobody at your shop ever picks up. Your insurer satisfaction scores took the hit, the front counter turned it off, and now every time somebody says the words "AI phone" in your building, the service advisor makes a face.
Fair. The 2024 version could not see your repair orders, could not tell a first-notice call from a status call, and cheerfully read out a promise date that had changed two days earlier. What is different in 2026 is not only that these agents talk like people and answer inside about two-tenths of a second. It is that you can now grade one against your own recorded calls before it ever speaks to a customer, watch what it did step by step on every single test case, and hand it authority in stages instead of all at once.
It failed for reasons that have nothing to do with speech. A body shop's phone is not a restaurant's phone. Roughly two-thirds of your inbound calls are people asking about a car that is already in your building, and the right answer depends on a repair order status that changes hourly — in paint, waiting on parts, waiting on approval, ready for QC. A phone robot that cannot read the RO is a robot that can only take a message, and taking a message is the thing that made the customer angry in the first place.
Grading an agent means measuring it against real past calls where you already know what the right answer was, watching each step it took to get there, and only widening what it is allowed to do once it passes. That is the discipline that matured in 2026, and it is the reason this conversation is different from the one you had two years ago.
Your phone system already records calls, or can be turned on to. Export six months. In most shops that is a few thousand recordings, and you are not listening to all of them — you are pulling roughly 200 that cover the shape of your week. Include the Monday 8am rush, the 4:30 to 5:30 pickup window and every Saturday you are open.
Sort them into the buckets your counter actually sees: status and promise date; first notice of loss and tow-in; "do you work with my insurance"; price questions on a bumper or a door; parts delay complaints; rental car questions; scan and calibration questions, usually starting with "what is this $350 line"; and the angry ones — paint match, a comeback, a total-loss dispute. Add a ninth if it fits your shop: vendors, adjusters, glass subcontractors and tow operators calling in.
This is the work, and it is a couple of afternoons. For each of the 200 calls, write the answer that was correct on the day: the RO number, the status in the system at that hour, whether the part was on national backorder, whether the customer had authorised the supplement, what your CSR actually said and whether that was right. Where the CSR got it wrong, write what should have happened. That set of cases is your test, and nobody else can build it for you — it is made from your ROs, your insurers, your promise dates.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for auto shop in your browser — 60 seconds, no signup.
Then write the rules the agent must never break. Mine would be: never quote a repair price over the phone, ever, for any damage, on any car. Always capture insurer, claim number, vehicle and a callback number. Hand off immediately on the words attorney, injury, lawsuit, total loss dispute, or any request for the owner. Only read a promise date if the RO was touched in the last 24 hours; otherwise say a person will call back with a firm date. Never book a drop-off into a day that is already at capacity in the blueprint stall.
flowchart TD
A["Export six months of front-counter recordings"] --> B["Sort into eight call types"]
B --> C["Status and promise-date calls"]
B --> D["First notice of loss and tow-in calls"]
B --> E["Insurance, rental and calibration questions"]
C --> F["Score against what the RO actually said that day"]
D --> F
E --> F
F --> G{"Captures the claim, quotes no prices, hands off cleanly?"}
G -->|No| B
G -->|Yes| H["Live after 5pm only, still watched daily"]
A transcript tells you what the agent said. It does not tell you whether it said the right thing for the right reason, and that difference is what bites you three months in. What you want to see for every test call is the trail: which repair order it opened, which field it read, when it last refreshed, what it decided, what it wrote back.
An agent that tells a customer "your vehicle is in refinish" because it read the RO status is fine. An agent that says the same sentence because the caller mentioned paint and it guessed is a problem waiting for a Saturday morning. Both produce identical transcripts. Only the step-by-step trail separates them, and being able to see that trail is exactly what got usable in 2026.
Stage one, listen only. The agent runs alongside your CSR on live calls and answers nothing. You compare what it would have said against what your person said. Two weeks.
Stage two, after hours only. It answers from 5:30pm to 7:30am and on Sundays, takes messages, captures the claim and vehicle, and books nothing. Every call gets reviewed the next morning by your office manager. Two weeks.
Stage three, read the RO. It may state status and promise date under the 24-hour freshness rule, and it may text the customer a link. Still no bookings, still reviewed daily.
Stage four, book. It schedules estimate appointments and drop-offs into the capacity your production manager has set, and only then. If your CSI survey scores or your comeback rate move the wrong way at any stage, you go back a stage. Nobody has to have an argument about it, because you set the rule before you started.
Illustration only. Use your own call logs.
Still reading? Stop comparing — try CallSphere live.
See the auto shop AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
| Inbound calls per month | 640 |
| Unanswered or abandoned | 18 percent, about 115 |
| Of those, new-work inquiries | 15 percent, about 17 |
| Convert to a written estimate | 40 percent, about 7 |
| Estimates that become repair orders | 50 percent, about 3.4 |
| Average repair order | $4,300 |
| Gross profit at 42 percent | about $6,140 a month |
Track two other things alongside the money, because they are what your DRP scorecard reflects: how many status calls got a correct answer without reaching a person, and whether your customer satisfaction scores moved. If the money is up and the satisfaction score is down, you are trading next year's referrals for this month's revenue and you should stop.
The comeback call. Somebody stands in your lot pointing at a door that does not match in daylight, and they are not calling for information — they are calling to find out whether you are the kind of shop that takes care of it. That is your general manager, on the phone, that day.
The total-loss call. When a customer learns their car is not coming back, they need somebody who can walk them through what the carrier does next, what happens to the personal belongings in the trunk, and how the storage bill works. Hand it off on the first sentence.
And anything with injury or a lawyer in it. If the caller mentions being hurt, or an attorney, or the other driver's carrier disputing liability, the agent takes a name and a number and gets a person on the line. Write that rule down and test it explicitly, because it is the one that costs the most when it is missed.
Two afternoons for 200 calls if your office manager does it with the repair order history open beside them. It feels like a lot of work for a phone system. It is the only part of this that decides whether it works, and you keep it — every new agent you consider gets graded against the same 200 calls.
Most will work it out, and it should say so if asked. What moves your satisfaction scores is not whether it is a machine — it is whether the call got answered and the answer was right. A customer who gets a correct promise date at 7:15pm is more forgiving than one who reaches a voicemail box at 4:45 in the afternoon.
Volume triples and your counter drowns, which is when an agent earns its keep. Build the test set including a CAT week if you have one recorded, and set a rule for it: during a declared event the agent captures name, vehicle, insurer and claim, books nothing, and tells the caller when the shop will call back. Overpromising during a storm is how shops lose customers for two years.
Usually not. In most shops the agent sits in front of the existing numbers and hands calls through to the counter, so the physical phones on the desk keep working exactly as they do now.
CallSphere builds the AI voice and chat agents this post is about — answering the shop line and the website chat 24 hours a day, capturing the insurer, claim and vehicle, booking estimate appointments and drop-offs, and passing the calls that need a person to a person. The part worth insisting on with any vendor, including us: bring your own 200 recorded calls, grade the agent on them, and do not widen what it is allowed to do until it earns it.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Vietnamese, Spanish and Mandarin-speaking dental front offices order the minimum and ask nothing. Live translation in 2026 changes the lunch-hour call.
The 200-call test set a treatment program should build from its own recordings, the five things to score, and how to widen an agent's authority safely.
How an industrial equipment OEM builds a parts-desk test set from its own closed orders, grades the agent on as-built revisions, and stops wrong-part shipments.
Community management agents fail on escalations and claims, not tone. Build a 400-thread test set from your own social inbox before anything goes live.
Cattle producers can now measure a phone agent against real past calls, watch every step it took, and widen what it is allowed to do only once it earns it.
How a design firm builds an RFI test set from closed jobs, scores citation and routing accuracy, and widens an agent's authority only when the numbers earn it.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI