By Sagar Shankaran, Founder of CallSphere
How a travel agency owner builds a graded test set from last wave season's client threads, scores it bucket by bucket, and widens the agent's authority safely.
Key takeaways
Here is the question, and it is not rhetorical: would you let a machine answer, with nobody watching, the 9 p.m. Sunday message that says "Hi — is my final payment due this Friday or next? And can I still add my mother to the cabin?"
Most owners I ask say no, then admit they have no way to find out. They watched a demo where an agent answered five questions beautifully. They have no idea what it does with the sixth. So the agent gets bolted onto the website as a glorified contact form, answers nothing that matters, and everybody concludes AI is not ready for travel. That is a testing problem, and in 2026 it got solved.
In this trade a wrong answer is not an embarrassment. It is a priced liability, and there are only four families of it.
The first is a money deadline: telling someone final payment is next Friday when the cruise line's date was this Friday, the booking auto-cancels, the cabin category is gone, and you eat the difference. The second is documents and entry: saying a passport is fine when it expires four months after travel and the destination wants six, or that the name on the ticket "should be close enough" when Secure Flight wants the government ID name. The third is refunds: any sentence starting "you'll definitely get your money back" about a supplier penalty schedule you do not control. The fourth is insurance and medical — any answer at all about whether a policy covers a pre-existing condition or a hurricane. That fourth one is where the errors-and-omissions policy stops being theoretical.
An agent evaluation is nothing more than a graded exam you write from your own past customer messages, where you already know the right answer, run before the agent is allowed to talk to anybody live. It is unglamorous, which is why almost nobody did it before the tooling made it easy.
The 2024 version of this was a transcript. You read what the agent said, decided it sounded fine, and turned it on. When it went wrong three weeks later you had no idea why, so your fix was another paragraph of instructions and hope.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for hospitality in your browser — 60 seconds, no signup.
Through 2025 and into 2026, the evaluation and monitoring tooling around these agents matured into the thing that gates a rollout. You run the agent against a batch of real past cases in one go, and for each you get a step-by-step record: which of your documents it opened, which supplier's penalty schedule it read, whether it checked the booking record or guessed, where it stopped and handed off. When it gives a wrong final-payment date, you can see it pulled the Alaska sailing's terms while answering about the Caribbean one. That is a fixable fault. "It said something weird" is not.
flowchart TD
A["Pull 200 real inquiries from ClientBase and the after-hours voicemail"] --> B["Sort into five buckets, write the correct answer for each"]
B --> C["Run the agent through all 200 with nobody helping it"]
C --> D["Read the step-by-step record of what it opened and checked"]
D --> E{"Did this bucket pass at 98 percent or better?"}
E -->|No| F["Fix the approved documents, rerun only that bucket"]
F --> C
E -->|Yes| G["Agent answers that bucket live, transcripts read daily"]
G --> H["Widen to the next bucket after two clean weeks"]
You already own the exam. Pull the last full wave season — 1 January through 31 March — because that is when your inquiry mix is richest and most stressful. Export threads from ClientBase Online or TravelJoy, the Viator and GetYourGuide message centers if you sell tours there, the shared inbox, and the after-hours duty phone voicemails. Take 200. Do not curate them to be nice.
Then sort into five buckets and, for each case, have the person who handled it write the correct answer — with the source. That last part is the work: roughly a day and a half of a senior reservations agent's time, and the highest-value day and a half you will spend on AI this year.
Suppose the agent runs your 200 and passes 176 — 88 percent. That sounds good and it is dangerous, because the failures are not spread evenly. Illustrative assumptions, using an agency fielding about 6,000 inquiries a year across phone, chat and the message centers.
| Bucket | Cases | Passed | Share of annual volume | Verdict |
|---|---|---|---|---|
| 1 — Shape and availability | 70 | 69 (99%) | 44% | Go live |
| 2 — Money deadlines | 45 | 36 (80%) | 19% | Hold; read-only lookups first |
| 3 — Documents and entry | 40 | 33 (83%) | 17% | Hold; fix source documents |
| 4 — Disruption | 30 | 24 (80%) | 13% | Hand off to human, always |
| 5 — Insurance and medical | 15 | 14 (93%) | 7% | Hand off only; never answer |
Turn that into money. Loose on everything at 88 percent, roughly 720 wrong answers a year go out. Assume one in twenty of the Bucket 2 mistakes ends in a cancelled booking rather than an apology. Bucket 2 is 19 percent of 6,000, so 1,140 inquiries; at 20 percent failure that is 228 wrong answers, and one in twenty of those is about 11 blown bookings. At an average $6,400 booking and 12 percent commission, that is roughly $8,800 of commission gone, before goodwill credits and the client who never comes back.
Now the restricted version: it answers Bucket 1 only. That is 44 percent of 6,000 — about 2,640 inquiries a year at 99 percent, risk cost close to zero. If each would otherwise have eaten four minutes of a reservations agent's time, you freed roughly 176 hours, which at $34 fully loaded is about $6,000. And you have a written record of exactly what to fix in Buckets 2 and 3. The exam cost a day and a half of senior time, call it $500.
Some things belong on a permanent no list. I would put four there.
Anything that reads as advice about an insurance policy. Not a summary, not "generally these policies cover," not a link with a sentence of interpretation. Coverage questions go to a licensed person or the insurer's own line, and the agent should say so and transfer.
Still reading? Stop comparing — try CallSphere live.
See the hospitality AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
Anything involving a medical condition, mobility need or accessibility accommodation. A guest mentioning a wheelchair, an oxygen concentrator or a dialysis schedule is telling you something operations must own with the supplier. The agent's job is to capture it and get a human on it inside the hour.
Anything that moves money — taking a card over chat, changing a payment plan, applying a future cruise credit. Your chargeback ratio in this trade is too fragile, and a disputed travel charge is expensive to fight even when you win.
And anything about a supplier in trouble. If an operator is wobbling or a carrier cancelled a season, the answer comes from you personally, with your Seller of Travel obligations in mind — a California CST or Florida ST registration means your name is on the disclosure. Not a place for a machine to improvise.
Do not shop for software on Monday. Export 200 threads from last wave season and have your most senior reservations agent write correct answers to 40 of them by Friday, noting where each came from — which cruise line's terms page, which operator's penalty schedule, which State Department page. Do the rest the following week. When a vendor shows up, hand them your 200 and ask for the step-by-step record on the failures. A vendor who does not want your exam has told you everything.
Build your own for anything client-facing: the liability sits with whoever gave the answer, and your client mix is not the host's average. Where the host helps is supplier terms — point the agent at their maintained penalty schedules rather than your own copies, which go stale by May.
Rerun the whole set whenever you change the agent's instructions or source documents, and once before each wave season. Add every real-world failure as it happens; a set that never grows stops catching anything by year two.
For low-stakes buckets, 98 percent, and read every transcript daily for a fortnight. For anything touching money, deadlines or documents, do not go live on the agent answering at all — go live on it drafting for a human to send. Most of the time saving, none of the liability.
Both, and phone matters more, because your after-hours line is the one place a wrong answer gets no second reader. Build the phone exam from voicemails and call recordings from the same wave season, score the same five buckets, and add one check: did it recognise an upset caller and get out of the way.
The phone and the chat window are where bookings are won in this trade, and both ring hardest when nobody is at a desk. CallSphere builds AI voice and chat agents that answer the line and website chat 24/7, book appointments and capture the inquiry. Bring one in the way described above: give it the bucket it can prove it handles, keep the rest on a human, and widen only when your 200 cases say so.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Set AI ceilings around wave season instead of the calendar month, give the air desk the strongest model, and judge the whole line item on cost per booking.
Before an AI answers your PT clinic's intake line, build a 150-call answer key from last year's schedule. What to sample, how to score, where humans stay.
How a community paper builds a 120-call test set from its own obituary intakes, runs an AI agent in shadow mode, and widens its authority one rung at a time.
Casino surveillance footage cannot leave the property. On-premises AI in 2026 cuts a 90-minute disputed handpay review to nine, with no video sent to a vendor.
The pre-season exterior survey eats four weeks of May and never gets finished. What autonomous inspection routes change for coastal vacation rental managers.
What an unplanned coach breakdown costs a tour operator, the engine readings that warn you first, and how to test condition-based monitoring on two coaches.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI