---
title: "A Tenant Calls at 9:52 PM About a Flooded Restroom. Grade the Agent on 200 Old Work Orders First."
description: "Build a 200-case test set from your own janitorial work orders, grade the agent, and widen its authority in stages. Includes the five pass/fail call types."
canonical: https://callsphere.ai/blog/a-tenant-calls-at-9-52-pm-about-a-flooded-restroom-grade-the-agent-on-
category: "Home Services"
tags: ["commercial cleaning", "janitorial", "after-hours calls", "ai agents", "quality control", "work orders"]
author: "CallSphere Team"
published: 2026-07-10T13:54:06.000Z
updated: 2026-09-09T00:39:54.758Z
---

# A Tenant Calls at 9:52 PM About a Flooded Restroom. Grade the Agent on 200 Old Work Orders First.

> Build a 200-case test set from your own janitorial work orders, grade the agent, and widen its authority in stages. Includes the five pass/fail call types.

It is 9:52 on a Tuesday night. A tenant on the third floor of a suburban office building calls the number on the sticker by the elevator because there is water coming out from under the men's room door and across the corridor carpet. Your answering service picks up, takes a message, and emails it to a shared inbox that your account manager will open at 7:15 the next morning. By then the carpet has been wet for nine hours, the property manager has already sent his own email, and the phrase "failure to perform" has appeared in writing.

Every owner in this trade knows that sequence. The obvious fix in 2026 is an agent that answers that line, understands what it just heard, and does something. The less obvious part — the part that separates the companies where this works from the companies where it blows up in month two — is what you do before you point that agent at a real customer.

## What the after-hours line costs you today

Price out the current arrangement honestly. The answering service is a few hundred dollars a month, which is cheap, and it buys you a message. What it does not buy is triage. A flooded restroom, a request for a carpet extraction quote, a complaint about trash not pulled on the second floor, and a locked-out tenant all arrive as the same yellow message slip.

That flattening is where the money leaks. The flood needed a callback crew inside an hour. The carpet extraction was a billable periodic worth $1,400 that nobody quoted for eleven days. The trash complaint was the second one that month on an account with a 30-day cure clause. And the account manager who reads all four at 7:15 AM handles them in the order they appear in the inbox, which has nothing to do with the order they should be handled in.

The clean way to say what changed in 2026: you can now watch, step by step, exactly what an agent did with a real customer call, replay it against calls that already happened, and score it before you let it speak to anyone new.

## The 2026 change: evaluation became the gate, not the afterthought

In 2024 you bought a voice agent, turned it on, and hoped. There was no honest way to know whether it was right, so the first evidence you got was an angry property manager. This year the tooling that measures agents against real cases and shows every step it took matured into the thing that gates a rollout. That is a boring sentence with a large consequence: the sensible order of operations is now measure first, widen authority second.

What you are buying is not just an agent. You are buying the record: for every call, what it heard, how it classified it, which of your documents it consulted, what it told the caller, and what it did in your work order system. When it gets one wrong, you can see where it turned left.

```mermaid
flowchart TD
  A["Tenant calls the service line at 9:52 PM"] --> B["Agent classifies against your own past work orders"]
  B --> C["Emergency: water, sewage or biohazard"]
  B --> D["Next-shift add-on for the night lead"]
  B --> E["Billable extra — carpet extraction quote"]
  B --> F["Complaint that starts the 30-day cure clock"]
  C --> G["Account manager paged, callback crew dispatched"]
  D --> G
  E --> G
  F --> G
  G --> H["Every step logged with the recording and transcript"]
```

## Build the test set out of two hundred calls you already have

This is the work, and it is one long afternoon, not a project. Pull two hundred real items from the last twelve months. Your sources are already sitting there: the work order log in CleanTelligent, Janitorial Manager, Lighthouse or whatever inspection software you run; the eHub or WinTeam service request history; and the honest one, the account manager's sent folder.

Pick a mix that reflects your actual portfolio, not the exciting ones. If 60% of after-hours calls are trash and restroom supply complaints, 60% of your test set should be trash and restroom supply complaints. Include the awkward items on purpose: the caller who is not a tenant but a vendor locked in the loading dock; the one who called about a smell in the parking garage that was not yours; the property manager who called to say a cleaner left a stripper bottle on a windowsill.

For each one, write down the right answer, decided by you and your operations manager, in three fields: which category it is, what should happen tonight, and who gets told. That written sheet is the test. Everything else is opinion.

## The five answers that must never be wrong

Grade everything, but treat five categories as pass/fail, because getting them wrong costs real money the same week.

- **Standing water, sewage, or blood.** These are not "we will handle it on the next shift." Bloodborne pathogen work has its own OSHA rule and your crew needs the right people and the right kit. The agent must page a human, every time.
- **Anything out of scope.** If a tenant asks for the interior glass done or the break room refrigerator emptied and that is not in the specification, the agent must not agree. It logs it as a change order for the account manager to price. One "sure, we will take care of that" becomes a year of free work.
- **Promising a callback.** A callback crew at 10:30 PM is overtime, drive time, and sometimes a second person for security reasons. That decision has a dollar figure and needs a named approver.
- **Keys, codes and alarm panels.** The agent never says a code, never confirms who holds a key, never discusses which cleaner is assigned to which suite.
- **The second complaint on the same account in thirty days.** This is the one owners miss. It is not a service call, it is a contract event, and it must reach you, not just the account manager.

## Scoring, then widening authority in three steps

Assumptions, illustrative: 200 graded cases, and you run them before anyone is live. Say the agent classifies 182 correctly, or 91%. That number alone tells you nothing useful. The breakdown does.

| Miss type | Count | What it would have cost |
| --- | --- | --- |
| Billable extra treated as a complaint | 7 | Quote never sent — say $1,400 each |
| Trash complaint escalated as emergency | 6 | Needless callback, ~$180 each |
| Out-of-scope request accepted | 3 | Free work, and a precedent |
| Standing water routed to next shift | 2 | Unacceptable — blocks go-live |

Those two water misses are the whole story. Ninety-one percent is a passing grade in school and a failing grade here, because the two failures are the two that produce a claim. Fix the routing rule for water, run the two hundred again, and only then go live — first in listen-and-draft mode where it writes the message and a human sends it, then live on the two safest categories only, then wider.

## Where the human stays, permanently

Some of this should never be handed over, no matter how good the scores get. An angry property manager who is deciding whether to put your contract out to bid gets your operations manager on the phone, not an agent. A cleaner calling in an injury gets a human, immediately, because that call is the start of a workers' compensation file and a possible OSHA recordable.

And the quarterly business review stays human, obviously. The agent can produce the inspection summary and the trend on complaints by building, which is genuinely useful and saves your account manager a Sunday. It cannot sit across a table from a facilities director and read the room.

## Frequently asked questions

### How long does it really take to build the test set?

Pulling the two hundred items takes about two hours if your inspection software exports cleanly. Writing the correct answers takes four to six hours with two people, and it must be two people, because the arguments are the valuable part. Budget one working day and stop treating it as a technology task — it is a supervision task.

### What if my complaints do not live in software at all?

Then they live in an inbox, and that is fine. Search the account manager's mail for the words your customers actually use — flood, no trash, restroom out, smell, missed. Two hundred is not a large number in a year of email for a 30-building portfolio. If you truly cannot find two hundred, use one hundred and go live more slowly.

### Can the agent create the work order itself?

Eventually, and that is the point of doing this in stages. Start with it drafting the work order and a human clicking save. When you have several weeks where nobody had to correct the draft, let it save for the two safest categories. Keep water, biohazard and anything out of scope on a human hand indefinitely.

### Will the property manager know they are talking to an agent?

Tell them. Say it in the first sentence of the call and put it in the account transition letter. Some states have their own disclosure rules and more are coming, but the practical reason is simpler: a property manager who feels tricked at 10 PM will remember it at renewal.

If you decide the after-hours line is worth fixing, [CallSphere](https://callsphere.ai) builds AI voice and chat agents that answer business phone lines and web chat around the clock, route the call, book the walk-through and capture the lead with the building and contact attached. Do it in the order this post describes: score it against your own two hundred calls first, keep the water and the out-of-scope requests on a human, and widen only after the record shows it earned it.

---

Source: https://callsphere.ai/blog/a-tenant-calls-at-9-52-pm-about-a-flooded-restroom-grade-the-agent-on-
