You Tried a Phone Bot in 2024 and It Promised Scallops You Didn't Have. Build the 80-Call Test Set First.
By Sagar Shankaran, Founder of CallSphere
Grade an AI order-line agent against 80 real calls, credit memos and closure notices before it talks to a buyer. What to score, and what it must always refuse.
Key takeaways
The 2024 phone bot failed because nobody could see what it had said
You already tried this. Some time in 2024 a vendor put a bot on the order line, and inside a week it told a Boston chef you had 400 pounds of U-10 dry scallops ready for Thursday. You had ninety pounds, all wet-pack, and the box truck driver found out at the loading door. The office manager pulled the plug that afternoon and nobody in the building has brought it up since.
That reaction was correct. The 2024 version had no way of knowing what was in the cooler, and — this mattered more — you had no way of checking what it had been telling people until a customer complained.
What changed in 2026 is not that the agent learned about scallop grades. It is that the tooling for testing and watching one grew up. Through this year the practical question shifted from "does it sound good in a demo" to "show me, case by case, what it did." You hand it a batch of saved situations, it works each one, and you get back what it decided, what record it opened, and the exact words it produced. You compare that against the answer you wrote down in advance. Anything that breaks a hard rule shows up as a red line, not as a customer complaint three weeks later.
The other half is authority. You do not hand an agent the order line on day one. You give it the narrowest job it can do — telling callers whether a boat has landed and taking a callback number — and widen the job only after it has cleared the exam twice in a row. The exam comes first. The phone comes later.
A test set is nothing more exotic than a folder of calls your business already handled, each one paired with the answer you would have accepted, run against the agent before a single customer hears it.
What the order desk actually handles between 4 a.m. and 9 a.m.
At a lobster and groundfish dealer with two trucks and a 12,000-pound holding system, the first five hours of the day are the whole business. The boats call in with ETAs and rough counts. A Portland restaurant group wants today's board price on selects and whether you have any hard-shell left in the tank. A distributor's buyer wants forty pounds of scallops added to a standing Thursday order and wants to know the count. A chef asks whether the oysters he bought last week were out of an area that closed after Sunday's rain. A new account asks whether you can invoice net-30 and whether your shipper number is current.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent in your browser — 60 seconds, no signup.
Every one of those is a different kind of risk. The price question is a margin decision. The availability question is a promise. The harvest area question is a food-safety question with an inspector attached to it. And the person best qualified to answer all four is standing on the dock in bibs watching a hoist, not sitting at the desk.
So the workaround is the answering machine, the callback, and the thing everyone pretends is fine: the buyer calls a competitor while waiting. Nobody counts those. They do not show up in any report, because a call that was never returned leaves no record.
flowchart TD
A["Pull 80 real calls from last season"] --> B["Write the answer you would have accepted"]
B --> C["Run the agent against all 80 offline"]
C --> D{"Did it break a hard rule?"}
D -->|Yes| E["Fix the rule and re-run the same 80"]
E --> C
D -->|No| F["Shadow week: agent drafts, order clerk sends"]
F --> G["Widen to price and availability calls before 6 a.m."]
Where the 80 cases come from in a seafood house
You do not write them. You already have them, in four places most dealers never think of as records.
- The credit memo file. This is the single richest source in the building. Every credit is a case where somebody was told the wrong thing — short weight, wrong count, wrong pack, arrived warm. Take the last thirty.
- The order inbox. Twelve months of emails from chefs and distributors, including the messy ones: "same as last week but swap the 10-20s for U-10 if you have them, otherwise call me."
- The text threads with the boats. These teach the agent what a real landing message looks like, including the ones with no punctuation sent from a wheelhouse at 3 a.m.
- The closure notices. Every state shellfish closure and reopening notice from last season, so you can test whether the agent refuses to sell out of a closed area.
Aim for a spread: roughly forty routine orders, twenty-five awkward ones, and fifteen the agent must refuse outright. The refusals matter most. If it will not promise scallops that are not in the cooler count, and will not confirm a shipment from a rainfall-closed area, it has avoided the two failures that cost you an account.
The grading rules a seafood order desk actually needs
Write your rules in the language you would use with a new hire on their first Monday. Half of them are hard fails — one violation and the run is a failure, no averaging.
- Never quote a price that is not on today's board. If the board has not been posted, take a number and say the office will call back.
- Never promise a species, grade or count that is not in the last cooler count. "I think we have" is a failure.
- Never confirm product from a harvest area listed as closed or conditionally approved and currently in closure.
- Always capture the lot or tag reference on any shellfish order line.
- Hand off immediately on a weight dispute, a credit request, a quality complaint, or anything a health department asks.
- Say it is an automated assistant at the start of the call, every time.
That last one is not a nicety. Under Texas TRAIGA and California SB 53, both in force since 1 January 2026, and under the EU AI Act transparency obligations carrying a 2 August 2026 compliance date, disclosure that a person is dealing with an AI system is the baseline expectation. If you ship live lobster to buyers in Spain or Italy, the European rules can reach you. Put the disclosure in the first sentence and stop thinking about it.
What the exam costs, in office hours
Assumptions, all illustrative for a two-truck dealer: the office manager earns roughly $34 an hour fully loaded; it takes about six minutes to pull a case and write the accepted answer; you run three rounds of fixes; and one blown promise on a wholesale account costs you a week of that account's business.
| Item | Basis | Cost |
|---|---|---|
| Building the 80-case set | 80 cases × 6 min = 8 hours | $272 |
| Three fix-and-re-run rounds | 3 × 90 min review | $153 |
| Shadow week (clerk reviews drafts) | 5 days × 40 min | $113 |
| Total to get to a gated rollout | $538 | |
| One blown promise on a $6,000-a-week account | One week lost, no return | $6,000 |
The exam pays for itself if it prevents one bad promise a decade. It will prevent more than that in the first month, because the refusal cases are the ones the old bot got wrong.
Still reading? Stop comparing — try CallSphere live.
CallSphere ships complete AI voice agents per industry — 14 tools for healthcare, 10 agents for real estate, 4 specialists for salons. See how it actually handles a call before you book a demo.
Where the agent should never get the last word
Price on a soft market day is a judgement call about the relationship, not a lookup. When the boats are landing hard and the board is falling, whether you hold a number for a five-year account is the owner's decision, and it should stay one.
Weight disputes stay human. So does anything touching a recall or a health department inquiry — those go straight to the person who signs your food-safety records, with the call transcript attached. And the boats stay human. A captain calling in with a fouled hydraulic and a hold full of product does not want an assistant; he wants the person who can move a truck.
The agent's honest job is the routine calls between 4 a.m. and 7 a.m. that nobody is free to answer, and the after-hours calls that reach a machine.
Frequently asked questions
How many cases is enough to be worth doing?
Eighty is a good working number for a dealer with one order line, and you can start meaningfully at forty. What matters more than the count is the mix. Twenty routine orders will teach you nothing you did not know. Fifteen situations where the correct answer is "no, and here is why" will tell you within an afternoon whether the agent is safe to put near a buyer.
We do not record our calls. Can we still do this?
Yes, and this is the common case. Build the set from written traffic instead: the order inbox, the text threads, the credit memos, and the closure notices. Then have whoever answers the phone spend two weeks jotting one line about every unusual call on a legal pad by the desk. That pad becomes the second half of your test set, and it costs nobody an extra hour.
What happens when the cooler count is wrong?
Then the agent is wrong, confidently, in your voice. This is the failure mode to plan for. Give it a rule that anything more than a set number of hours old is treated as unknown, and have it say "let me have the dock confirm that and call you right back" rather than reading a stale figure. An agent that says "I need to check" is worth ten that guess.
Do we have to tell buyers it is not a person?
Say it in the first sentence. Beyond the state and European rules already in force in 2026, buyers dislike finding out afterwards, and discovering it mid-negotiation costs more goodwill than the disclosure ever will.
Start Monday with the credit memo folder
Pull the last thirty credit memos and write, next to each one, the sentence somebody should have said. That is three hours of work and it is the hardest part of the whole exercise. Everything after it is running the same thirty cases again and again until the agent stops failing them.
At CallSphere we build the voice and chat agents that sit on business phone lines and web chat — answering after hours, taking orders and callbacks, and booking the call back with a human. The reason we bring up the test set first is that a seafood order desk is not a forgiving place to learn: the promise you cannot fill is remembered a lot longer than the call you answered on the first ring. Test it against your own history, watch what it does, then widen it.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
Try CallSphere AI Voice Agents
See how AI voice agents work for your industry. Live demo available -- no signup required.