---
title: "Grade the Ticket Agent on 300 Closed ConnectWise Tickets Before It Answers One Live"
description: "Build an AI help desk test set from your own ConnectWise or Autotask history, score it by category, and find the security tickets it must never close."
canonical: https://callsphere.ai/blog/grade-the-ticket-agent-on-300-closed-connectwise-tickets-before-it-ans
category: "IT & SaaS Support"
tags: ["managed service provider", "help desk automation", "connectwise psa", "ai evaluation", "msp operations", "tier 1 support"]
author: "CallSphere Team"
published: 2026-06-09T07:48:12.000Z
updated: 2026-07-25T23:21:34.198Z
---

# Grade the Ticket Agent on 300 Closed ConnectWise Tickets Before It Answers One Live

> Build an AI help desk test set from your own ConnectWise or Autotask history, score it by category, and find the security tickets it must never close.

Would you hand a brand-new Tier 1 hire the service board on their first morning, point at the queue, and walk away? Nobody does. You sit them next to a senior tech, you read every ticket they close for two weeks, and you keep the ones that smell wrong away from them entirely. Then somebody shows you an AI help desk agent, runs a canned demo where it resets a password in eleven seconds, and asks you to point it at your ConnectWise service board on Monday.

The demo proves nothing. What changed in 2026 is that you no longer have to take the vendor's word for it. Agent evaluation and step-by-step review tooling matured to the point where you can grade an agent the way you would grade a new hire — against work you have already done, with the right answer known in advance — before it touches a single live client.

## What a help desk agent actually gets wrong

The failures that worry a managed service provider owner are not the ones the vendor demos. They are these:

- It answers a question for the wrong client. Two of your customers have a user named J. Martinez. One is on a HIPAA business associate agreement, one is a 22-person insurance agency, and the agent reads the wrong documentation page.
- It resolves the surface complaint and buries the real one. "Outlook keeps asking me for my password" gets a profile rebuild and a closed ticket, when the sign-in logs in the Microsoft 365 admin center show the session was revoked because somebody in Lagos had a stolen sign-in cookie.
- It closes a ticket that should have started an SLA clock. The contact was the controller at your largest account, calling at 4:52 p.m. on the last business day of the month.
- It writes a time entry that is plausible and wrong, which quietly bleeds into your agreement true-up and your effective hourly rate.

Every one of those is invisible in a demo and obvious in your own ticket history. Which is exactly where the test set comes from.

## Build the test set out of your own closed tickets

**An evaluation set, in help desk terms, is a graded pile of your own already-closed tickets that you run the agent against before it answers a live one — you already know the correct outcome for each, so you can score the agent instead of trusting it.**

Pull the last 300 closed tickets from ConnectWise PSA, Autotask, or HaloPSA. Do not pull them at random and do not pull the easy ones. Build the pile deliberately:

- 120 genuinely routine: password and multi-factor resets, mapped drive gone, printer offline, mailbox permission, VPN client won't connect, "my laptop is slow after the Windows update."
- 60 that look routine and are not: the mailbox rule that forwards to an outside address, the "new laptop for a new hire" that was actually a termination, the shared mailbox request from someone who should not have it.
- 40 cross-client traps: same first name, same software, two different documentation records in IT Glue or Hudu.
- 40 tied to a specific contract term: after-hours coverage, on-site only, a client whose agreement excludes third-party line-of-business software.
- 40 that ended in escalation to your security operations partner — Huntress, Blackpoint, whoever holds the pager.

Then have your service manager write the right answer on each one: correct action, correct board, correct escalation, correct client. That labeling is a day of somebody's life. It is the cheapest day you will spend this year.

```mermaid
flowchart TD
  A["Pull last 300 closed tickets from ConnectWise"] --> B["Service manager writes the right answer on each"]
  B --> C["Agent runs all 300 in a sandbox, no live access"]
  C --> D{"Escalated all 9 security tickets?"}
  D -->|No| E["Fix the rules, re-run the same 300"]
  E --> C
  D -->|Yes| F{"Wrong action under 4 percent?"}
  F -->|No| E
  F -->|Yes| G["Go live on password and Outlook tickets only"]
  G --> H["Dispatcher reads every closed ticket for two weeks"]
```

## The nine tickets that decide it, and watching how it got there

In a pile of 300 real tickets from a 900-device shop, you will typically find a handful that started as "can't sign in" and ended as an incident. Those are your gate. Not a soft target — a gate. If the agent resolves even one of them as a password problem, it does not go live, no matter how good the other 291 look.

This is the part owners get backwards. They set an overall accuracy target — "95% and we ship it" — and the 5% it misses is exactly the 5% that costs you a client. Score by category, not in aggregate. Routine tickets can be graded on speed and correctness. Security-adjacent tickets are graded pass/fail on one question: did it stop and hand off to a human?

The other half of what matured in 2026 is the ability to see what the agent did on the way to its answer, ticket by ticket: which documentation record it opened, which client it believed it was working for, which admin action it proposed, where it hesitated. That review screen is what turns a bad outcome into a fixable one.

When your agent closes a printer ticket wrongly, you want to see that it read the Hudu page for the client's old print server — the one your project team decommissioned in March and nobody archived. That is not an AI problem. That is a documentation hygiene problem the agent just found for you, and half the shops that run this exercise report the same thing: the test run audits your documentation as a side effect.

## The arithmetic on a 4% miss rate

Illustration, not a case study. Assume a shop with 1,400 tickets a month across 42 clients, a loaded Tier 1 cost of $52 an hour, and an average of 11 minutes of technician time on the routine tickets the agent would handle.

| Line | Assumption | Result |
| --- | --- | --- |
| Tickets per month | Across 42 clients | 1,400 |
| Share the agent is allowed to touch | Password, Outlook profile, printer, drive mapping only | 38% = 532 |
| Resolved correctly | 91% of in-scope | 484 |
| Escalated correctly | 6% | 32 |
| Wrong action | 3% | 16 |
| Technician minutes returned | 484 × 11 min | 88.7 hours |
| Rework on the 16 misses | 25 min each | 6.7 hours |
| Net hours back per month | 88.7 &minus; 6.7 | 82 hours |
| Value at $52/hour loaded | 82 × $52 | $4,264/month |

Now the other side of the ledger. One missed compromise that reaches a client's finance mailbox costs you an incident response engagement, roughly a week of your best engineer, a very unpleasant call with the client's cyber insurance carrier, and a real chance of losing the account at renewal. That single number swamps the $4,264. Which is why the gate is not "is it accurate enough on average" — it is "does it refuse to touch the dangerous ones."

## Where to keep a human, permanently

Some categories should never come off the human board, no matter how well the agent scores:

- Anything that changes a permission — mailbox delegation, admin role, shared drive access, conditional access exceptions.
- Anything involving a departing employee. Terminations are contentious, legally sensitive, and frequently mis-described by the person calling.
- Anything from a client under a compliance obligation you are attesting to — CMMC Level 2, HIPAA, PCI — where the audit trail matters more than the fix.
- Any ticket where the contact is angry. That is a relationship event, not a technical one, and it belongs to your account manager.

Write those four exclusions into the agent's boundaries before you write anything else, and put them in the test set as tickets it is supposed to hand off.

## What to do Monday

Export 300 closed tickets. Put your service manager on the labeling for one day. Run the agent against them in a sandbox with no write access to your PSA at all — nothing it does can create, close, or bill anything. Score by category. If it clears the security gate and lands under 4% wrong actions, turn it on for two ticket types only and have your dispatcher read every closed ticket for two weeks. Widen it a category at a time, re-running the same 300 each time you change something.

## Frequently asked questions

### How many tickets do I actually need in the test set?

Three hundred is enough to catch category-level failures for a shop under about 1,500 tickets a month. The number matters less than the mix. A hundred well-chosen tickets that include your genuinely nasty cases will tell you more than a thousand password resets.

### Won't the agent just memorize the test tickets?

It can, which is why you hold back a second set of 100 tickets it never sees during tuning and only run those at the end. Same discipline you would use with any certification exam.

### My PSA data is a mess. Do I have to clean it first?

No, and cleaning it first is a trap that kills the project. Run the evaluation on the messy data, because messy data is what the agent will face on Tuesday. The failures will tell you exactly which documentation records and ticket types to fix, in priority order, which is far better than a blanket cleanup.

### Should I tell clients that an agent is answering their tickets?

Yes, before it happens, in writing, in a one-paragraph addendum to the agreement. It costs you nothing when they hear it from you and it costs you a client when they figure it out on their own. Several state AI statutes that took effect on 1 January 2026 also push toward disclosure, and if any of your clients touch EU users, the transparency obligations with a 2 August 2026 compliance date are worth a conversation with your attorney.

## A note on the phones

The ticket board is only half your intake. The other half rings, usually while every technician is already on a call, and the same evaluation discipline applies: score a voice agent against your own recorded calls and your own after-hours log before it answers a client. [CallSphere](https://callsphere.ai) builds AI voice and chat agents that answer business phone lines and web chat, take the details, book the appointment, and capture the lead around the clock — and the sane way to bring one in is the same as above: prove it on the calls you have already handled, then widen its authority one call type at a time.

---

Source: https://callsphere.ai/blog/grade-the-ticket-agent-on-300-closed-connectwise-tickets-before-it-ans
