By Sagar Shankaran, Founder of CallSphere
Grade an AI order-line agent against 80 real calls, credit memos and closure notices before it talks to a buyer. What to score, and what it must always refuse.
Key takeaways
You already tried this. Some time in 2024 a vendor put a bot on the order line, and inside a week it told a Boston chef you had 400 pounds of U-10 dry scallops ready for Thursday. You had ninety pounds, all wet-pack, and the box truck driver found out at the loading door. The office manager pulled the plug that afternoon and nobody in the building has brought it up since.
That reaction was correct. The 2024 version had no way of knowing what was in the cooler, and — this mattered more — you had no way of checking what it had been telling people until a customer complained.
What changed in 2026 is not that the agent learned about scallop grades. It is that the tooling for testing and watching one grew up. Through this year the practical question shifted from "does it sound good in a demo" to "show me, case by case, what it did." You hand it a batch of saved situations, it works each one, and you get back what it decided, what record it opened, and the exact words it produced. You compare that against the answer you wrote down in advance. Anything that breaks a hard rule shows up as a red line, not as a customer complaint three weeks later.
The other half is authority. You do not hand an agent the order line on day one. You give it the narrowest job it can do — telling callers whether a boat has landed and taking a callback number — and widen the job only after it has cleared the exam twice in a row. The exam comes first. The phone comes later.
A test set is nothing more exotic than a folder of calls your business already handled, each one paired with the answer you would have accepted, run against the agent before a single customer hears it.
At a lobster and groundfish dealer with two trucks and a 12,000-pound holding system, the first five hours of the day are the whole business. The boats call in with ETAs and rough counts. A Portland restaurant group wants today's board price on selects and whether you have any hard-shell left in the tank. A distributor's buyer wants forty pounds of scallops added to a standing Thursday order and wants to know the count. A chef asks whether the oysters he bought last week were out of an area that closed after Sunday's rain. A new account asks whether you can invoice net-30 and whether your shipper number is current.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent in your browser — 60 seconds, no signup.
Every one of those is a different kind of risk. The price question is a margin decision. The availability question is a promise. The harvest area question is a food-safety question with an inspector attached to it. And the person best qualified to answer all four is standing on the dock in bibs watching a hoist, not sitting at the desk.
So the workaround is the answering machine, the callback, and the thing everyone pretends is fine: the buyer calls a competitor while waiting. Nobody counts those. They do not show up in any report, because a call that was never returned leaves no record.
flowchart TD
A["Pull 80 real calls from last season"] --> B["Write the answer you would have accepted"]
B --> C["Run the agent against all 80 offline"]
C --> D{"Did it break a hard rule?"}
D -->|Yes| E["Fix the rule and re-run the same 80"]
E --> C
D -->|No| F["Shadow week: agent drafts, order clerk sends"]
F --> G["Widen to price and availability calls before 6 a.m."]
You do not write them. You already have them, in four places most dealers never think of as records.
Aim for a spread: roughly forty routine orders, twenty-five awkward ones, and fifteen the agent must refuse outright. The refusals matter most. If it will not promise scallops that are not in the cooler count, and will not confirm a shipment from a rainfall-closed area, it has avoided the two failures that cost you an account.
Write your rules in the language you would use with a new hire on their first Monday. Half of them are hard fails — one violation and the run is a failure, no averaging.
That last one is not a nicety. Under Texas TRAIGA and California SB 53, both in force since 1 January 2026, and under the EU AI Act transparency obligations carrying a 2 August 2026 compliance date, disclosure that a person is dealing with an AI system is the baseline expectation. If you ship live lobster to buyers in Spain or Italy, the European rules can reach you. Put the disclosure in the first sentence and stop thinking about it.
Assumptions, all illustrative for a two-truck dealer: the office manager earns roughly $34 an hour fully loaded; it takes about six minutes to pull a case and write the accepted answer; you run three rounds of fixes; and one blown promise on a wholesale account costs you a week of that account's business.
| Item | Basis | Cost |
|---|---|---|
| Building the 80-case set | 80 cases × 6 min = 8 hours | $272 |
| Three fix-and-re-run rounds | 3 × 90 min review | $153 |
| Shadow week (clerk reviews drafts) | 5 days × 40 min | $113 |
| Total to get to a gated rollout | $538 | |
| One blown promise on a $6,000-a-week account | One week lost, no return | $6,000 |
The exam pays for itself if it prevents one bad promise a decade. It will prevent more than that in the first month, because the refusal cases are the ones the old bot got wrong.
Still reading? Stop comparing — try CallSphere live.
CallSphere ships complete AI voice agents per industry — 14 tools for healthcare, 10 agents for real estate, 4 specialists for salons. See how it actually handles a call before you book a demo.
Price on a soft market day is a judgement call about the relationship, not a lookup. When the boats are landing hard and the board is falling, whether you hold a number for a five-year account is the owner's decision, and it should stay one.
Weight disputes stay human. So does anything touching a recall or a health department inquiry — those go straight to the person who signs your food-safety records, with the call transcript attached. And the boats stay human. A captain calling in with a fouled hydraulic and a hold full of product does not want an assistant; he wants the person who can move a truck.
The agent's honest job is the routine calls between 4 a.m. and 7 a.m. that nobody is free to answer, and the after-hours calls that reach a machine.
Eighty is a good working number for a dealer with one order line, and you can start meaningfully at forty. What matters more than the count is the mix. Twenty routine orders will teach you nothing you did not know. Fifteen situations where the correct answer is "no, and here is why" will tell you within an afternoon whether the agent is safe to put near a buyer.
Yes, and this is the common case. Build the set from written traffic instead: the order inbox, the text threads, the credit memos, and the closure notices. Then have whoever answers the phone spend two weeks jotting one line about every unusual call on a legal pad by the desk. That pad becomes the second half of your test set, and it costs nobody an extra hour.
Then the agent is wrong, confidently, in your voice. This is the failure mode to plan for. Give it a rule that anything more than a set number of hours old is treated as unknown, and have it say "let me have the dock confirm that and call you right back" rather than reading a stale figure. An agent that says "I need to check" is worth ten that guess.
Say it in the first sentence. Beyond the state and European rules already in force in 2026, buyers dislike finding out afterwards, and discovering it mid-negotiation costs more goodwill than the disclosure ever will.
Pull the last thirty credit memos and write, next to each one, the sentence somebody should have said. That is three hours of work and it is the hardest part of the whole exercise. Everything after it is running the same thirty cases again and again until the agent stops failing them.
At CallSphere we build the voice and chat agents that sit on business phone lines and web chat — answering after hours, taking orders and callbacks, and booking the call back with a human. The reason we bring up the test set first is that a seafood order desk is not a forgiving place to learn: the promise you cannot fill is remembered a lot longer than the call you answered on the first ring. Test it against your own history, watch what it does, then widen it.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Vietnamese, Spanish and Mandarin-speaking dental front offices order the minimum and ask nothing. Live translation in 2026 changes the lunch-hour call.
The 200-call test set a treatment program should build from its own recordings, the five things to score, and how to widen an agent's authority safely.
Sort six months of front-counter recordings into eight call types, write the right answer for each, then widen the agent's authority in four stages, not one.
Abandoned calls in the cutoff hour cost broadline distributors real gross profit. What a 200-millisecond voice agent on the order desk changes, with the math.
No integration exists for eLandings or dealer reporting. A 2026 AI agent drives the screens itself, files the landing and holds anything it cannot read.
How a solar installer builds a test set from past service calls, watches the agent step by step, and widens its authority in stages without repeating 2024.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI