By Sagar Shankaran, Founder of CallSphere
Score a drive-thru voice agent against orders your own crew already rang up: the five numbers, the accuracy floors, and how to widen its authority safely.
Key takeaways
You already tried this. Somewhere around 2023 a vendor put voice ordering on the outside order point at your busiest store. It asked "anything else?" four times, it could not hear a woman with a toddler in the back seat, the crew hit the override button on roughly half the cars, and you pulled it out after nine weeks with a bad taste and an invoice.
Fair. Here is what is genuinely different in 2026, and it is not the voice. The voice was always going to get better — speech-to-speech answers now come back in about two-tenths of a second, which is faster than a crew member who is also bagging, and the newer speech recognition handles a running engine and a car radio far better than the 2023 generation did. That part is solved enough.
What changed is that you can now prove whether it works on your menu, at your store, before a guest ever hears it.
Think about how that pilot was judged. The vendor demonstrated it on a quiet Tuesday at 3:00 p.m. with the regional manager standing in the lane. You went live at lunch. The crew complained. The GM said it was slowing the line. You looked at the drive-thru timer report, saw the average order time was up eleven seconds, and killed it.
Nobody ever established what the crew's own accuracy was, on the same orders, at the same hour. Nobody counted which items it got wrong. Nobody could go back and see why it heard "no pickle" as "more pickle" on car 41. It was a vibe, judged against another vibe.
Through the first half of 2026 the tooling for exactly that — measuring an agent against real cases, then watching, step by step, what it actually did on each one — matured into the thing that gates a rollout. In serious operations nobody widens an agent's authority now without it.
Agent evaluation means scoring the software against orders your own crew already rang up, before a single guest ever speaks to it.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for restaurant in your browser — 60 seconds, no signup.
You do not have to invent a test. Your stores generate one every day, and if you run a drive-thru with recording, you already have the audio and the matching check from Toast, Aloha or PAR Brink sitting in the same minute of the same day.
Pull 300 of them, deliberately, not at random. Fifty from the 12:05–12:35 peak, when three cars are queued and the order taker is also on the headset for the window. Twenty-five from 7:40 a.m., when the order is coffee and the guest mumbles. Forty with heavy modifiers — no onion, extra pickle, sub the side, dressing on the side, make one of them a combo with the Freestyle machine. Twenty in Spanish. Twenty where a second person in the car changes the order at the total. Fifteen with a coupon code or an app offer. Twenty on the limited-time item that launched three weeks ago, because that is where a new agent fails hardest — it either does not know the item exists or rings it at the wrong price. Ten with an allergen request, which the agent must never answer and must always hand to a person. And twenty of the ugly ones: a wrong-lane car, a DoorDash driver at the speaker, a guest asking whether the lobby is open.
flowchart TD
A["Pull 300 recorded lunch-rush orders with the ticket the crew rang"] --> B["Agent takes each order offline, no guest involved"]
B --> C{"Does the agent ticket match what the crew actually rang?"}
C -->|Below your accuracy floor| D["Fix menu wording, modifiers and prices, run the 300 again"]
D --> B
C -->|At or above floor| E["Live on overflow phone orders only, GM reviews 20 a day"]
E --> F["Widen to drive-thru tickets under $25"]
F --> G["Catering and large group orders stay with a person"]
Score item-level accuracy, not order-level. "Order was right" hides too much; a six-item family order with one wrong side is a remake either way, but you need to know whether it is always the same side.
Then measure four more. Time from the guest's first word to the total appearing on the board, because that is the number the drive-thru timer is going to judge you on at 12:15. Escalation rate — how often it hands to a person, which should be high in week one and should never be zero. Allergen handling, which is pass or fail with no middle ground: any mention of an allergy hands off, every time, or you do not go live. And attach rate, because if it never asks about the drink you have quietly given away your best upsell.
Set the floors before you look at the results, in writing, with your Director of Operations. If you set them afterwards you will move them, because you will have already spent the money.
Illustration with stated assumptions. One store, 1,000 drive-thru tickets a week. Assume your crew currently rings 2.6% of them with a wrong or missing item — use your own remake and comp data, not this figure. Assume each error costs $3.40 in remade food and $9.80 in comps and goodwill, so $13.20 all in.
| Line | Crew today | Agent tests at 1.9% | Agent tests at 3.9% |
|---|---|---|---|
| Wrong tickets per week | 26 | 19 | 39 |
| Cost per week at $13.20 | $343 | $251 | $515 |
| Cost per year, one store | $17,850 | $13,050 | $26,780 |
| Versus today | — | saves $4,800 | costs $8,930 |
Now the honest part about 300 tickets. A test set that size will tell you very clearly whether the agent is at two percent or six percent. It cannot reliably tell you two percent from two-point-six, because at those rates you are talking about the difference between six errors and eight in the sample. So use the 300 to catch the disasters, which it does brilliantly, and get the fine number from the first fortnight of live orders with every ticket reviewed.
The rollout order that works in this trade is not the one vendors propose. Start where a failure is cheap and invisible: overflow phone orders that currently ring out during peak, where the alternative is nobody answering at all. Two weeks there, with the GM listening to twenty a day.
Then drive-thru tickets under $25 at one store, at off-peak dayparts only, with the crew override button on the headset and a standing instruction that any crew member may take a car at any time with no explanation required. Then peak. Then the second store. Catering and large group orders never move — a $780 office order has a name, a delivery window and a person who will phone you personally, and that is worth a human every time.
Still reading? Stop comparing — try CallSphere live.
See the restaurant AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
Watch what it actually did, order by order, in the first weeks. Not a score. The transcript, the item it rang, the moment it decided the guest said "large". That is where you find out that your menu names two things "chicken bowl" and always has, and that is why car 41 got the wrong one.
Allergens, permanently. Not because the agent cannot read your allergen matrix, but because the liability of a wrong answer at a speaker box is not a liability you should ever hand to software, and because your guest deserves to hear a person say it.
The car that changes its mind three times, the guest who asks for the manager, the confused older regular who orders the same thing every Thursday and expects to be recognised, the lane during a POS outage. And the crew override stays, forever. The day you take away a shift lead's ability to grab a car is the day the whole thing becomes something the crew is fighting rather than using.
One more limit worth stating plainly: this does not fix a slow kitchen. If your average order-to-window time is bad because the make line is understaffed at 12:15, a faster order point just moves the queue to a different place.
Monday's first step: pull fifty recorded orders and the matching checks, and read them yourself. Before you evaluate any agent, you will learn something uncomfortable and useful about how your own crew takes an order at peak.
Realistically two days of somebody's time to pull, label and match 300 orders, plus a few hours a quarter to keep it current as the menu changes. It is the single highest-return administrative task in this whole project, and it is also the one everybody wants to skip.
Assume yes, and check your state. Several US states have AI disclosure rules on the books now, with more taking effect through 2026, and the penalties are cheaper to avoid than to argue about. A short line at the start of the greeting is fine and, in practice, most guests stop noticing after the first visit.
They will hate it if it makes their shift harder, and they will defend it if it takes the headset off them during a rush so they can bag. Which of those you get is decided by whether the override button works and whether you asked the shift leads what to fix after week one.
Then this belongs on your phone line rather than your speaker post, and that is often the better first move anyway. Ask your Franchise Business Consultant in writing which order channels you control as a franchisee. The answer is frequently that the drive-thru is the brand's and the restaurant phone is yours.
That phone line is the easiest place to start, because the failure mode today is an unanswered ring during peak. CallSphere builds AI voice and chat agents that answer the restaurant phone and web chat around the clock, take the catering enquiry, book the callback and capture the lead — and everything above about building a test set from your own recordings applies to that line first.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
How an industrial equipment OEM builds a parts-desk test set from its own closed orders, grades the agent on as-built revisions, and stops wrong-part shipments.
Community management agents fail on escalations and claims, not tone. Build a 400-thread test set from your own social inbox before anything goes live.
Cattle producers can now measure a phone agent against real past calls, watch every step it took, and widen what it is allowed to do only once it earns it.
Sort six months of front-counter recordings into eight call types, write the right answer for each, then widen the agent's authority in four stages, not one.
DoorDash adjustments, short cases and off-contract invoice lines quietly cost six stores real money. What a 2026 overnight agent has finished before 6 am.
How a design firm builds an RFI test set from closed jobs, scores citation and routing accuracy, and widens an agent's authority only when the numbers earn it.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI