By Sagar Shankaran, Founder of CallSphere
Build an AI help desk test set from your own ConnectWise or Autotask history, score it by category, and find the security tickets it must never close.
Key takeaways
Would you hand a brand-new Tier 1 hire the service board on their first morning, point at the queue, and walk away? Nobody does. You sit them next to a senior tech, you read every ticket they close for two weeks, and you keep the ones that smell wrong away from them entirely. Then somebody shows you an AI help desk agent, runs a canned demo where it resets a password in eleven seconds, and asks you to point it at your ConnectWise service board on Monday.
The demo proves nothing. What changed in 2026 is that you no longer have to take the vendor's word for it. Agent evaluation and step-by-step review tooling matured to the point where you can grade an agent the way you would grade a new hire — against work you have already done, with the right answer known in advance — before it touches a single live client.
The failures that worry a managed service provider owner are not the ones the vendor demos. They are these:
Every one of those is invisible in a demo and obvious in your own ticket history. Which is exactly where the test set comes from.
An evaluation set, in help desk terms, is a graded pile of your own already-closed tickets that you run the agent against before it answers a live one — you already know the correct outcome for each, so you can score the agent instead of trusting it.
Pull the last 300 closed tickets from ConnectWise PSA, Autotask, or HaloPSA. Do not pull them at random and do not pull the easy ones. Build the pile deliberately:
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for IT support in your browser — 60 seconds, no signup.
Then have your service manager write the right answer on each one: correct action, correct board, correct escalation, correct client. That labeling is a day of somebody's life. It is the cheapest day you will spend this year.
flowchart TD
A["Pull last 300 closed tickets from ConnectWise"] --> B["Service manager writes the right answer on each"]
B --> C["Agent runs all 300 in a sandbox, no live access"]
C --> D{"Escalated all 9 security tickets?"}
D -->|No| E["Fix the rules, re-run the same 300"]
E --> C
D -->|Yes| F{"Wrong action under 4 percent?"}
F -->|No| E
F -->|Yes| G["Go live on password and Outlook tickets only"]
G --> H["Dispatcher reads every closed ticket for two weeks"]
In a pile of 300 real tickets from a 900-device shop, you will typically find a handful that started as "can't sign in" and ended as an incident. Those are your gate. Not a soft target — a gate. If the agent resolves even one of them as a password problem, it does not go live, no matter how good the other 291 look.
This is the part owners get backwards. They set an overall accuracy target — "95% and we ship it" — and the 5% it misses is exactly the 5% that costs you a client. Score by category, not in aggregate. Routine tickets can be graded on speed and correctness. Security-adjacent tickets are graded pass/fail on one question: did it stop and hand off to a human?
The other half of what matured in 2026 is the ability to see what the agent did on the way to its answer, ticket by ticket: which documentation record it opened, which client it believed it was working for, which admin action it proposed, where it hesitated. That review screen is what turns a bad outcome into a fixable one.
When your agent closes a printer ticket wrongly, you want to see that it read the Hudu page for the client's old print server — the one your project team decommissioned in March and nobody archived. That is not an AI problem. That is a documentation hygiene problem the agent just found for you, and half the shops that run this exercise report the same thing: the test run audits your documentation as a side effect.
Illustration, not a case study. Assume a shop with 1,400 tickets a month across 42 clients, a loaded Tier 1 cost of $52 an hour, and an average of 11 minutes of technician time on the routine tickets the agent would handle.
| Line | Assumption | Result |
|---|---|---|
| Tickets per month | Across 42 clients | 1,400 |
| Share the agent is allowed to touch | Password, Outlook profile, printer, drive mapping only | 38% = 532 |
| Resolved correctly | 91% of in-scope | 484 |
| Escalated correctly | 6% | 32 |
| Wrong action | 3% | 16 |
| Technician minutes returned | 484 × 11 min | 88.7 hours |
| Rework on the 16 misses | 25 min each | 6.7 hours |
| Net hours back per month | 88.7 − 6.7 | 82 hours |
| Value at $52/hour loaded | 82 × $52 | $4,264/month |
Now the other side of the ledger. One missed compromise that reaches a client's finance mailbox costs you an incident response engagement, roughly a week of your best engineer, a very unpleasant call with the client's cyber insurance carrier, and a real chance of losing the account at renewal. That single number swamps the $4,264. Which is why the gate is not "is it accurate enough on average" — it is "does it refuse to touch the dangerous ones."
Still reading? Stop comparing — try CallSphere live.
See the IT support AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
Some categories should never come off the human board, no matter how well the agent scores:
Write those four exclusions into the agent's boundaries before you write anything else, and put them in the test set as tickets it is supposed to hand off.
Export 300 closed tickets. Put your service manager on the labeling for one day. Run the agent against them in a sandbox with no write access to your PSA at all — nothing it does can create, close, or bill anything. Score by category. If it clears the security gate and lands under 4% wrong actions, turn it on for two ticket types only and have your dispatcher read every closed ticket for two weeks. Widen it a category at a time, re-running the same 300 each time you change something.
Three hundred is enough to catch category-level failures for a shop under about 1,500 tickets a month. The number matters less than the mix. A hundred well-chosen tickets that include your genuinely nasty cases will tell you more than a thousand password resets.
It can, which is why you hold back a second set of 100 tickets it never sees during tuning and only run those at the end. Same discipline you would use with any certification exam.
No, and cleaning it first is a trap that kills the project. Run the evaluation on the messy data, because messy data is what the agent will face on Tuesday. The failures will tell you exactly which documentation records and ticket types to fix, in priority order, which is far better than a blanket cleanup.
Yes, before it happens, in writing, in a one-paragraph addendum to the agreement. It costs you nothing when they hear it from you and it costs you a client when they figure it out on their own. Several state AI statutes that took effect on 1 January 2026 also push toward disclosure, and if any of your clients touch EU users, the transparency obligations with a 2 August 2026 compliance date are worth a conversation with your attorney.
The ticket board is only half your intake. The other half rings, usually while every technician is already on a call, and the same evaluation discipline applies: score a voice agent against your own recorded calls and your own after-hours log before it answers a client. CallSphere builds AI voice and chat agents that answer business phone lines and web chat, take the details, book the appointment, and capture the lead around the clock — and the sane way to bring one in is the same as above: prove it on the calls you have already handled, then widen its authority one call type at a time.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
The 200-call test set a treatment program should build from its own recordings, the five things to score, and how to widen an agent's authority safely.
Shift leads, housekeepers and floor staff skip your help desk because of language. Gemini live translation puts them on the call with your on-call engineer.
Build an AI test set from closed mortgage files: the 1008 qualifying income is the answer key, QC defects are the hard cases, and the breakdown beats the score.
MSP receiving, warranty registration and closet documentation still run on hand-typed serials. 2026 readers pull them from phone photos, faxes and handwriting.
How a caterer builds a 240-case test set from real BEOs and inquiry email, what it costs to grade it, and the category floors to hit before going live.
MSP quarterly business reviews cost 4-6 vCIO hours per client. Claude Cowork and ChatGPT Work return the deck, summary and license variance sheet by morning.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI