By Sagar Shankaran, Founder of CallSphere
Build a 200-case test set from your own janitorial work orders, grade the agent, and widen its authority in stages. Includes the five pass/fail call types.
Key takeaways
It is 9:52 on a Tuesday night. A tenant on the third floor of a suburban office building calls the number on the sticker by the elevator because there is water coming out from under the men's room door and across the corridor carpet. Your answering service picks up, takes a message, and emails it to a shared inbox that your account manager will open at 7:15 the next morning. By then the carpet has been wet for nine hours, the property manager has already sent his own email, and the phrase "failure to perform" has appeared in writing.
Every owner in this trade knows that sequence. The obvious fix in 2026 is an agent that answers that line, understands what it just heard, and does something. The less obvious part — the part that separates the companies where this works from the companies where it blows up in month two — is what you do before you point that agent at a real customer.
Price out the current arrangement honestly. The answering service is a few hundred dollars a month, which is cheap, and it buys you a message. What it does not buy is triage. A flooded restroom, a request for a carpet extraction quote, a complaint about trash not pulled on the second floor, and a locked-out tenant all arrive as the same yellow message slip.
That flattening is where the money leaks. The flood needed a callback crew inside an hour. The carpet extraction was a billable periodic worth $1,400 that nobody quoted for eleven days. The trash complaint was the second one that month on an account with a 30-day cure clause. And the account manager who reads all four at 7:15 AM handles them in the order they appear in the inbox, which has nothing to do with the order they should be handled in.
The clean way to say what changed in 2026: you can now watch, step by step, exactly what an agent did with a real customer call, replay it against calls that already happened, and score it before you let it speak to anyone new.
In 2024 you bought a voice agent, turned it on, and hoped. There was no honest way to know whether it was right, so the first evidence you got was an angry property manager. This year the tooling that measures agents against real cases and shows every step it took matured into the thing that gates a rollout. That is a boring sentence with a large consequence: the sensible order of operations is now measure first, widen authority second.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for home services in your browser — 60 seconds, no signup.
What you are buying is not just an agent. You are buying the record: for every call, what it heard, how it classified it, which of your documents it consulted, what it told the caller, and what it did in your work order system. When it gets one wrong, you can see where it turned left.
flowchart TD
A["Tenant calls the service line at 9:52 PM"] --> B["Agent classifies against your own past work orders"]
B --> C["Emergency: water, sewage or biohazard"]
B --> D["Next-shift add-on for the night lead"]
B --> E["Billable extra — carpet extraction quote"]
B --> F["Complaint that starts the 30-day cure clock"]
C --> G["Account manager paged, callback crew dispatched"]
D --> G
E --> G
F --> G
G --> H["Every step logged with the recording and transcript"]
This is the work, and it is one long afternoon, not a project. Pull two hundred real items from the last twelve months. Your sources are already sitting there: the work order log in CleanTelligent, Janitorial Manager, Lighthouse or whatever inspection software you run; the eHub or WinTeam service request history; and the honest one, the account manager's sent folder.
Pick a mix that reflects your actual portfolio, not the exciting ones. If 60% of after-hours calls are trash and restroom supply complaints, 60% of your test set should be trash and restroom supply complaints. Include the awkward items on purpose: the caller who is not a tenant but a vendor locked in the loading dock; the one who called about a smell in the parking garage that was not yours; the property manager who called to say a cleaner left a stripper bottle on a windowsill.
For each one, write down the right answer, decided by you and your operations manager, in three fields: which category it is, what should happen tonight, and who gets told. That written sheet is the test. Everything else is opinion.
Grade everything, but treat five categories as pass/fail, because getting them wrong costs real money the same week.
Assumptions, illustrative: 200 graded cases, and you run them before anyone is live. Say the agent classifies 182 correctly, or 91%. That number alone tells you nothing useful. The breakdown does.
| Miss type | Count | What it would have cost |
|---|---|---|
| Billable extra treated as a complaint | 7 | Quote never sent — say $1,400 each |
| Trash complaint escalated as emergency | 6 | Needless callback, ~$180 each |
| Out-of-scope request accepted | 3 | Free work, and a precedent |
| Standing water routed to next shift | 2 | Unacceptable — blocks go-live |
Those two water misses are the whole story. Ninety-one percent is a passing grade in school and a failing grade here, because the two failures are the two that produce a claim. Fix the routing rule for water, run the two hundred again, and only then go live — first in listen-and-draft mode where it writes the message and a human sends it, then live on the two safest categories only, then wider.
Still reading? Stop comparing — try CallSphere live.
See the home services AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
Some of this should never be handed over, no matter how good the scores get. An angry property manager who is deciding whether to put your contract out to bid gets your operations manager on the phone, not an agent. A cleaner calling in an injury gets a human, immediately, because that call is the start of a workers' compensation file and a possible OSHA recordable.
And the quarterly business review stays human, obviously. The agent can produce the inspection summary and the trend on complaints by building, which is genuinely useful and saves your account manager a Sunday. It cannot sit across a table from a facilities director and read the room.
Pulling the two hundred items takes about two hours if your inspection software exports cleanly. Writing the correct answers takes four to six hours with two people, and it must be two people, because the arguments are the valuable part. Budget one working day and stop treating it as a technology task — it is a supervision task.
Then they live in an inbox, and that is fine. Search the account manager's mail for the words your customers actually use — flood, no trash, restroom out, smell, missed. Two hundred is not a large number in a year of email for a 30-building portfolio. If you truly cannot find two hundred, use one hundred and go live more slowly.
Eventually, and that is the point of doing this in stages. Start with it drafting the work order and a human clicking save. When you have several weeks where nobody had to correct the draft, let it save for the two safest categories. Keep water, biohazard and anything out of scope on a human hand indefinitely.
Tell them. Say it in the first sentence of the call and put it in the account transition letter. Some states have their own disclosure rules and more are coming, but the practical reason is simpler: a property manager who feels tricked at 10 PM will remember it at renewal.
If you decide the after-hours line is worth fixing, CallSphere builds AI voice and chat agents that answer business phone lines and web chat around the clock, route the call, book the walk-through and capture the lead with the building and contact attached. Do it in the order this post describes: score it against your own two hundred calls first, keep the water and the out-of-scope requests on a human, and widen only after the record shows it earned it.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Follow one harvest-season day at a Havre, Montana family practice and see how an AI answering service catches the farm and ranch calls a busy desk cannot.
Williston, North Dakota urgent care clinics ride Bakken boom-bust call swings. An AI answering service scales up and down without adding front-desk payroll.
Minot, North Dakota pediatric clinics face winter call floods from base families and worried parents. An AI answering service books sick visits and routes fast.
Pediatric dental offices in Lincoln, Nebraska juggle busy parents and packed recall schedules. Seven phone fixes, starting with a 24/7 AI receptionist.
Dermatology clinics in Twin Falls, Idaho field skin-check calls from across the Magic Valley. An AI receptionist answers every one, 24/7, in two languages.
Cedar City, Utah urgent care clinics face festival-season call surges every summer. An AI answering service triages tourists and books visits around the clock.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI