By Sagar Shankaran, Founder of CallSphere
How a caterer builds a 240-case test set from real BEOs and inquiry email, what it costs to grade it, and the category floors to hit before going live.
Key takeaways
That objection comes up in about every third conversation, and it is fair. Plenty of caterers and venues put a chat widget on the website two years ago, watched it invent a Saturday in June, invent a per-person price, and tell somebody the salmon was gluten-free without anyone asking the chef. They pulled it down and have not thought about it since.
What changed is not that the machine got smarter about your business. It cannot be smart about your business; it has never seen your kitchen. What changed in 2026 is that the tooling to grade an agent before customers touch it grew up. You can now run it against a fixed list of real questions from your own history, see exactly what it did on each one and which document it pulled the answer from, and only then decide how much it is allowed to say. That gate is the product now.
An evaluation set is a fixed list of real questions taken from your own past events, each one paired with the answer you know is correct, that you run an AI agent against before it is ever allowed to speak to a customer.
Most owners hear "build a test set" and picture writing questions out of thin air. Do not. The questions already exist and they are sitting in three places you look at every week.
The first is your BEO archive — Caterease, Total Party Planner, Tripleseat, CaterZen, Better Cater, whichever one you run. Every banquet event order is a settled fact: this is what the client got, at this price, with these dietary notes, with these rentals, on this date. It is the closest thing your business has to a graded answer key.
The second is your inquiry email. Two years of threads between your catering sales manager and clients is two years of questions phrased the way real people phrase them — "do you allow an outside baker," "is the 22% a tip," "can we push the ceremony back if it rains," "what's the latest the band can play."
The third is the pile everyone forgets: the events that went sideways. The wedding where the count changed Thursday. The corporate lunch where nobody flagged the vegan. Those are the highest-value cases you own, because a wrong answer already cost you money once.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for hospitality in your browser — 60 seconds, no signup.
The difference between a 2024 chat widget and a 2026 agent is not the talking. It is that you can replay the whole thing. For any test question you see the answer, the step where it decided which document to look in, the document it actually used, and the moment it either stuck to your policy sheet or made something up. Wrong answers stop being mysterious. They sort into a few causes you can fix.
That is what turned "try it and see" into a rollout gate. Run the whole test set, get a score by category, look at the failures, fix the underlying document or narrow the agent's authority, re-run. Then — and only then — widen what it is allowed to answer. Nobody did this in 2024 because nobody could see inside the thing.
flowchart TD
A["Pull 240 real questions from BEOs and inquiry email"] --> B["Write the correct answer for each, chef signs the food ones"]
B --> C["Run the agent against all 240, read every failure"]
C --> D{"Zero wrong answers on allergens and contract terms?"}
D -->|No| E["Fix the policy sheet or block that question type"]
E --> C
D -->|Yes| F["Stage 1: answers 12 policy questions, live, recorded"]
F --> G{"Two weeks clean on real calls?"}
G -->|No| E
G -->|Yes| H["Stage 2: checks dates and books tours"]
Here is a workable breakdown for a venue doing weddings, corporate and social. The counts matter less than the categories.
Your Director of Catering writes the correct answer for the first three groups. Your executive chef signs off on the dietary group and on anything that names a dish. Do not let one person do both — the split is what makes the answer key trustworthy.
Building the answer key is real work and it is worth pricing it honestly.
| Item | Assumption | Cost |
| Writing correct answers | 240 cases × 6 minutes = 24 hours | 24 × $38/hr loaded = $912 |
| Chef sign-off on food cases | 3 hours × $46/hr loaded | $138 |
| Owner review of failures, two rounds | 4 hours | $220 |
| One-time total | $1,270 |
Now the other column. Suppose the agent quotes a plated dinner at last season's per-person price — $8 under current — on a 220-guest wedding, and you honor it because it came from your phone line. That single answer costs $1,760, which is more than the entire test set. A wrong allergen answer does not have a tidy number attached to it; it has an incident report, a comped event, an insurance call and a review you cannot delete.
Assume the agent lands around 89 percent on a first pass — 214 of 240 right. That is a normal first score and nowhere near good enough to talk to anyone. The useful part is the shape of the 26 failures: typically a dozen stale prices, eight policy documents that contradict each other, five dietary questions it should have refused, and one genuinely ambiguous case your own staff disagree on. Three of those four causes are your documents, not the agent.
Score three buckets, not two: right, wrong, and correctly refused. An agent that says "I'll have our chef call you about that within the hour" is scoring a point, not dodging. If you only score right and wrong, you will accidentally train yourself to reward an agent that guesses.
Set hard floors by category rather than one overall number. Zero tolerance on allergens, on anything that changes a signed contract, and on confirming a date. Ninety-five percent on written policy. Ninety on pricing, with the standing rule that any quote leaving the building is a draft until your sales manager sends it. An overall "92% accurate" figure hides the exact five failures that can hurt you.
Still reading? Stop comparing — try CallSphere live.
See the hospitality AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
And re-run the whole set on a schedule — after every menu price change, after you renegotiate with the rental house, and at minimum before wedding season and before the corporate holiday run. The test set is not a launch task. It is a seasonal one, like changing the menu.
The chef owns every word about food. Not as a policy preference — as the thing that keeps a nut-allergic guest out of an ambulance. The agent takes the name, guest count, specific allergy and event date; the chef calls back.
The Director of Catering owns anything with a signature under it. Guarantees, count changes, date moves, deposit schedules, cancellation terms. An agent can read a contract term aloud; it should not agree to a new one.
And the owner owns the exceptions. The corporate client who always gets the room at 10 percent off, the funeral luncheon that needs the minimum waived, the bride whose date moved twice for reasons you know about. Those are not test-set failures. They are judgment, and they are why a venue with a real relationship still beats a cheaper room down the road.
Two afternoons for the questions if you pull from real email rather than inventing them, plus a few hours from the chef. The slow part is not writing questions; it is discovering that your policy sheet, your website and your contract say three different things about corkage. Fixing that is the real work, and worth doing whether or not you ever turn an agent on.
It can be restricted to answer only from documents you approve, and to show you which one it used — which is exactly what you want. The catch is that your documents have to be right and current. Most operators find their published menu prices and their actual selling prices drifted apart sometime last season, and the agent surfaces that in about four minutes.
There is no single number, which is why category floors matter more. A practical gate: zero misses on allergens and contract terms, at least 95 percent on written policy, and every single one of the 20 edge cases either right or correctly handed to a person. Get there and let it answer twelve question types on live calls — not everything.
Re-run it, yes; rebuild it, no. That is the point of having the set written down. A re-run is an afternoon once the answer key exists, and it is the only way to know whether an update helped or quietly broke your corkage answer.
If the surface you are grading is your phone line and your website chat, CallSphere builds those agents for venues and caterers — answering inquiry calls, booking tours and tastings, and capturing lead details 24/7, with the food and contract questions routed to your chef and your Director of Catering rather than answered. Bring your own test set. Any vendor who will not let you run one before go-live is telling you something.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
The 200-call test set a treatment program should build from its own recordings, the five things to score, and how to widen an agent's authority safely.
The Thursday production packet - prep list, vendor POs, staffing, rentals - built as one goal. Worked food-waste math and the habits an owner must change.
Build an AI test set from closed mortgage files: the 1008 qualifying income is the answer key, QC defects are the hard cases, and the breakdown beats the score.
Build an AI help desk test set from your own ConnectWise or Autotask history, score it by category, and find the security tickets it must never close.
Most restaurant bookings come after close. See how a 2026 AI agent captures nights-and-weekends reservations and catering while you sleep.
Catering and event leads waste phone time. See how a 2026 AI agent qualifies inquiries 24/7 so you talk only to ready buyers.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI