---
title: "Wave Season Handed You 200 Real Inquiries. Make the AI Agent Pass Them Before It Answers One Live Traveler."
description: "How a travel agency owner builds a graded test set from last wave season's client threads, scores it bucket by bucket, and widens the agent's authority safely."
canonical: https://callsphere.ai/blog/wave-season-handed-you-200-real-inquiries-make-the-ai-agent-pass-them-
category: "Hotels & Hospitality"
tags: ["travel agencies", "ai agent evaluation", "travel advisor operations", "wave season", "customer inquiries"]
author: "CallSphere Team"
published: 2026-06-12T12:11:29.000Z
updated: 2026-07-25T23:12:38.715Z
---

# Wave Season Handed You 200 Real Inquiries. Make the AI Agent Pass Them Before It Answers One Live Traveler.

> How a travel agency owner builds a graded test set from last wave season's client threads, scores it bucket by bucket, and widens the agent's authority safely.

Here is the question, and it is not rhetorical: would you let a machine answer, with nobody watching, the 9 p.m. Sunday message that says *"Hi — is my final payment due this Friday or next? And can I still add my mother to the cabin?"*

Most owners I ask say no, then admit they have no way to find out. They watched a demo where an agent answered five questions beautifully. They have no idea what it does with the sixth. So the agent gets bolted onto the website as a glorified contact form, answers nothing that matters, and everybody concludes AI is not ready for travel. That is a testing problem, and in 2026 it got solved.

## The four answers that turn into an errors-and-omissions claim

In this trade a wrong answer is not an embarrassment. It is a priced liability, and there are only four families of it.

The first is a money deadline: telling someone final payment is next Friday when the cruise line's date was this Friday, the booking auto-cancels, the cabin category is gone, and you eat the difference. The second is documents and entry: saying a passport is fine when it expires four months after travel and the destination wants six, or that the name on the ticket "should be close enough" when Secure Flight wants the government ID name. The third is refunds: any sentence starting "you'll definitely get your money back" about a supplier penalty schedule you do not control. The fourth is insurance and medical — any answer at all about whether a policy covers a pre-existing condition or a hurricane. That fourth one is where the errors-and-omissions policy stops being theoretical.

**An agent evaluation is nothing more than a graded exam you write from your own past customer messages, where you already know the right answer, run before the agent is allowed to talk to anybody live.** It is unglamorous, which is why almost nobody did it before the tooling made it easy.

## What actually changed in 2026: you can watch it work, one step at a time

The 2024 version of this was a transcript. You read what the agent said, decided it sounded fine, and turned it on. When it went wrong three weeks later you had no idea why, so your fix was another paragraph of instructions and hope.

Through 2025 and into 2026, the evaluation and monitoring tooling around these agents matured into the thing that gates a rollout. You run the agent against a batch of real past cases in one go, and for each you get a step-by-step record: which of your documents it opened, which supplier's penalty schedule it read, whether it checked the booking record or guessed, where it stopped and handed off. When it gives a wrong final-payment date, you can see it pulled the Alaska sailing's terms while answering about the Caribbean one. That is a fixable fault. "It said something weird" is not.

```mermaid
flowchart TD
  A["Pull 200 real inquiries from ClientBase and the after-hours voicemail"] --> B["Sort into five buckets, write the correct answer for each"]
  B --> C["Run the agent through all 200 with nobody helping it"]
  C --> D["Read the step-by-step record of what it opened and checked"]
  D --> E{"Did this bucket pass at 98 percent or better?"}
  E -->|No| F["Fix the approved documents, rerun only that bucket"]
  F --> C
  E -->|Yes| G["Agent answers that bucket live, transcripts read daily"]
  G --> H["Widen to the next bucket after two clean weeks"]
```

## Building the test set from wave season: 200 threads, five buckets

You already own the exam. Pull the last full wave season — 1 January through 31 March — because that is when your inquiry mix is richest and most stressful. Export threads from ClientBase Online or TravelJoy, the Viator and GetYourGuide message centers if you sell tours there, the shared inbox, and the after-hours duty phone voicemails. Take 200. Do not curate them to be nice.

Then sort into five buckets and, for each case, have the person who handled it write the correct answer — with the source. That last part is the work: roughly a day and a half of a senior reservations agent's time, and the highest-value day and a half you will spend on AI this year.

- **Bucket 1 — Shape and availability.** "Do you have anything to Iceland in early March?" "Is the 7-night Alaska still open for four in a balcony?" Low risk, high volume.
- **Bucket 2 — Money deadlines.** Deposit amounts, final payment dates, penalty schedules, group name-change deadlines. High risk, and the answer lives in the supplier's terms, not in your head.
- **Bucket 3 — Documents and entry.** Passport validity, visas, ESTA, minors travelling with one parent, name-matching on air tickets, the Schengen 90-in-180 question that people get wrong every summer.
- **Bucket 4 — Disruption.** "My flight got cancelled, am I owed a refund?" "The hurricane is Tuesday — are we still going?" This bucket blew up your April.
- **Bucket 5 — Insurance and medical.** Every question about coverage. The correct answer here is almost always a warm hand-off to a licensed human, and that is what you grade against.

## The arithmetic: 200 cases, 24 failures, and what those 24 would have cost

Suppose the agent runs your 200 and passes 176 — 88 percent. That sounds good and it is dangerous, because the failures are not spread evenly. Illustrative assumptions, using an agency fielding about 6,000 inquiries a year across phone, chat and the message centers.

| Bucket | Cases | Passed | Share of annual volume | Verdict |
| --- | --- | --- | --- | --- |
| 1 — Shape and availability | 70 | 69 (99%) | 44% | Go live |
| 2 — Money deadlines | 45 | 36 (80%) | 19% | Hold; read-only lookups first |
| 3 — Documents and entry | 40 | 33 (83%) | 17% | Hold; fix source documents |
| 4 — Disruption | 30 | 24 (80%) | 13% | Hand off to human, always |
| 5 — Insurance and medical | 15 | 14 (93%) | 7% | Hand off only; never answer |

Turn that into money. Loose on everything at 88 percent, roughly 720 wrong answers a year go out. Assume one in twenty of the Bucket 2 mistakes ends in a cancelled booking rather than an apology. Bucket 2 is 19 percent of 6,000, so 1,140 inquiries; at 20 percent failure that is 228 wrong answers, and one in twenty of those is about 11 blown bookings. At an average $6,400 booking and 12 percent commission, that is roughly $8,800 of commission gone, before goodwill credits and the client who never comes back.

Now the restricted version: it answers Bucket 1 only. That is 44 percent of 6,000 — about 2,640 inquiries a year at 99 percent, risk cost close to zero. If each would otherwise have eaten four minutes of a reservations agent's time, you freed roughly 176 hours, which at $34 fully loaded is about $6,000. And you have a written record of exactly what to fix in Buckets 2 and 3. The exam cost a day and a half of senior time, call it $500.

## The questions that should never reach the agent at all

Some things belong on a permanent no list. I would put four there.

Anything that reads as advice about an insurance policy. Not a summary, not "generally these policies cover," not a link with a sentence of interpretation. Coverage questions go to a licensed person or the insurer's own line, and the agent should say so and transfer.

Anything involving a medical condition, mobility need or accessibility accommodation. A guest mentioning a wheelchair, an oxygen concentrator or a dialysis schedule is telling you something operations must own with the supplier. The agent's job is to capture it and get a human on it inside the hour.

Anything that moves money — taking a card over chat, changing a payment plan, applying a future cruise credit. Your chargeback ratio in this trade is too fragile, and a disputed travel charge is expensive to fight even when you win.

And anything about a supplier in trouble. If an operator is wobbling or a carrier cancelled a season, the answer comes from you personally, with your Seller of Travel obligations in mind — a California CST or Florida ST registration means your name is on the disclosure. Not a place for a machine to improvise.

## The Monday version of this

Do not shop for software on Monday. Export 200 threads from last wave season and have your most senior reservations agent write correct answers to 40 of them by Friday, noting where each came from — which cruise line's terms page, which operator's penalty schedule, which State Department page. Do the rest the following week. When a vendor shows up, hand them your 200 and ask for the step-by-step record on the failures. A vendor who does not want your exam has told you everything.

## Frequently asked questions

### We are an independent contractor under a host agency. Do we build our own test set, or does the host?

Build your own for anything client-facing: the liability sits with whoever gave the answer, and your client mix is not the host's average. Where the host helps is supplier terms — point the agent at their maintained penalty schedules rather than your own copies, which go stale by May.

### How often do I have to rerun the exam?

Rerun the whole set whenever you change the agent's instructions or source documents, and once before each wave season. Add every real-world failure as it happens; a set that never grows stops catching anything by year two.

### What score is good enough to go live on a bucket?

For low-stakes buckets, 98 percent, and read every transcript daily for a fortnight. For anything touching money, deadlines or documents, do not go live on the agent answering at all — go live on it *drafting* for a human to send. Most of the time saving, none of the liability.

### Does this apply to phone calls, or only to chat and email?

Both, and phone matters more, because your after-hours line is the one place a wrong answer gets no second reader. Build the phone exam from voicemails and call recordings from the same wave season, score the same five buckets, and add one check: did it recognise an upset caller and get out of the way.

The phone and the chat window are where bookings are won in this trade, and both ring hardest when nobody is at a desk. [CallSphere](https://callsphere.ai) builds AI voice and chat agents that answer the line and website chat 24/7, book appointments and capture the inquiry. Bring one in the way described above: give it the bucket it can prove it handles, keep the rest on a human, and widen only when your 200 cases say so.

---

Source: https://callsphere.ai/blog/wave-season-handed-you-200-real-inquiries-make-the-ai-agent-pass-them-
