---
title: "You Tried a Chatbot on 'My System Is Offline' Calls in 2024. Build the 200-Call Test Set First This Time."
description: "How a solar installer builds a test set from past service calls, watches the agent step by step, and widens its authority in stages without repeating 2024."
canonical: https://callsphere.ai/blog/you-tried-a-chatbot-on-my-system-is-offline-calls-in-2024-build-the-20
category: "Home Services"
tags: ["solar installers", "voice agents", "customer service", "agent evaluation", "service calls"]
author: "CallSphere Team"
published: 2026-06-17T14:01:19.000Z
updated: 2026-09-08T19:58:54.948Z
---

# You Tried a Chatbot on 'My System Is Offline' Calls in 2024. Build the 200-Call Test Set First This Time.

> How a solar installer builds a test set from past service calls, watches the agent step by step, and widens its authority in stages without repeating 2024.

## You tried this in 2024 and it embarrassed you

Most solar companies that put a chatbot on the service line two years ago took it off within a quarter. It told a customer her system was fine when the gateway had been offline for nine days. It answered a question about the federal tax credit that your office manager had to walk back on the phone. It offered a Tuesday appointment in a ZIP code where you have no crew. Somebody printed the transcript, it got passed around the office, and that was the end of that conversation for two years.

Fair. But the reason it failed was not that the machine could not talk. It was that nobody could see what it did between the customer's question and its answer, and nobody had ever tested it against real calls with known right answers before it went live. That is the thing that actually matured in 2026. Watching an agent step by step and grading it against your own past cases became ordinary tooling, and it is now the gate that a rollout has to pass rather than an afterthought.

**An agent evaluation for a solar installer is simply this: a folder of your own past customer calls with the correct answer written down for each one, run against the agent before it is allowed near a live line, and re-run every time anything changes.** If a vendor cannot show you that folder and the grades, they are asking you to repeat 2024.

## Four calls make up most of your inbound volume

Pull a month of your call log out of RingCentral or Aircall and tag them by hand. In a residential solar and storage company with a few thousand systems in the field, they cluster hard:

- **"My app says my system is offline."** Usually the gateway, usually the customer's new router, occasionally something real.
- **"My true-up bill was $1,340 and you told me my bill would be nothing."** A rate and expectations conversation, and the one most likely to be recorded and posted.
- **"Where are we?"** Anything between contract and permission to operate: permit status, install date, inspection, the utility.
- **"I am selling the house."** Lease and power purchase agreement transfers, plus the buyer's agent who wants a document by Thursday.

Every one of these has a correct answer that already exists somewhere — in your job tracker, in the monitoring portal, in the signed agreement, in the utility's application status. And every one of them has a wrong answer that costs real money. That combination is exactly what makes them worth testing rather than guessing at.

## Build the test set out of your own tickets before anything answers a live call

Take the last 200 inbound calls or web chats. Not made-up examples — the real ones, including the angry one and the one where the customer was confused about which company installed the system. For each, write down what the agent should have said, what it should have checked first, and what it should have refused to answer.

That writing is the work, and it takes an office manager about six minutes a call. It is also, incidentally, the best service audit you will ever run, because halfway through you will discover that three of your staff answer the true-up question three different ways.

```mermaid
flowchart TD
  A["Pull last 200 service calls from the ticket log"] --> B["Write the right answer for each one"]
  B --> C["Run the agent on all 200, read every step"]
  C --> D{"Did it ever guess at a bill, a guarantee or a credit?"}
  D -->|Yes| E["Narrow what it may say, run the 200 again"]
  E --> C
  D -->|No| F["Live on answering and intake only"]
  F --> G{"Two weeks clean on real calls?"}
  G -->|No| E
  G -->|Yes| H["Let it book the service visit"]
```

## Watch the steps, not just the answer

The final sentence the agent says is the least useful thing to grade. What you need to see is the sequence: did it identify the right site before it said anything about production, or did it match on a common last name? Did it actually read the monitoring status, or did it answer from the general shape of the question? When the customer said "my Powerwall," did it check whether that customer has storage at all, or did it play along?

That step-by-step record is what the 2026 tooling gives you and the 2024 chatbot did not. You can sit with your service manager and scroll through what the agent did on call 47, and she can point at the third step and say "that is where it went wrong — it should have asked for the service address first." Then you tighten that one rule and re-run all 200 in a few minutes rather than waiting for another customer to be the test.

Grade three things separately. **Correct**: it gave the answer your manager would have. **Safe but incomplete**: it did not know and handed off cleanly, which is a pass, not a failure. **Unsafe**: it stated something about money, coverage or equipment that it had no business stating. The unsafe count is the only one that has to be zero.

## Widen authority in three stages, not one

Stage one, read-only. It answers status questions and hands off everything else. Job status, install date, inspection scheduled, permission to operate received. No promises, no diagnosis, no dollars.

Stage two, intake and scheduling. It takes the service address, the system type, what the app is showing, whether there is a battery, and books a slot on a route day in the right region. It still refuses the money questions. This is where most of your saved hours are, because a full intake means your coordinator opens a ticket that is already complete instead of playing phone tag.

Stage three, narrow account actions, only if the earlier stages held for a month: sending the customer a copy of their agreement, confirming a lease transfer packet is on its way, resending a monitoring invitation. Each new power gets its own additions to the test set before it goes live. If you skip that step you are back to 2024, just with a better voice.

## The arithmetic on grading 200 calls

Assumptions, illustrative: 200 past calls, six minutes each to write the correct answer at a fully burdened $34 an hour, plus twelve hours of review time across two rounds of tightening. Live volume is 310 inbound calls a month.

| **Item** | **Amount** |
| --- | --- |
| Writing the answers: 20 hours | $680 |
| Two review rounds: 12 hours | $408 |
| One-time cost to build the test set | $1,088 |
| First run: unsafe answers | 9 of 200 (4.5%) |
| After tightening: unsafe answers | 0 of 200 |

Now price what you avoided. At a 4.5% unsafe rate on 310 calls a month, roughly 14 customers would have been told something wrong about a bill, a guarantee or their coverage. If one in five of those turns into a goodwill concession — a free service visit, a credit, a discounted cleaning to end the argument — at an average $900, that is about $2,520 a month you did not spend, against a one-time $1,088. And that is before the calls you miss entirely during install season.

## The sentences the agent must never be allowed to say

Write this list before you write anything else, and check it in the test set explicitly. No statement about what a customer's utility bill will be, ever — rate structures and true-up cycles are where trust dies. No statement about tax credit or incentive eligibility; that is between the customer and their accountant, and the rules moved recently enough that even your closers get it wrong. No promise about how long a battery will run a house in an outage. No determination of whether something is covered under workmanship warranty versus the manufacturer's warranty, because that decision has your money attached to it.

And absolutely nothing that instructs a customer to touch equipment. Not the DC disconnect, not the rapid shutdown switch, not the breaker in the combiner. "Go flip the switch on the side of the inverter" is how a homeowner ends up in front of energised conductors, and no amount of testing makes that an acceptable thing for an automated voice to say. The safe version is: we will have a technician call you back, and here is when.

Keep a human on the angry call, too. A customer who has been waiting eleven weeks for permission to operate while making loan payments does not want an efficient answer; they want somebody with authority to apologise and commit to a date. Route those to a person by keyword and by tone, and check in the test set that the routing actually fires.

## Frequently asked questions

### How long does building the test set actually take?

Two afternoons for the writing if your office manager blocks the time, plus a couple of hours per review round. The bottleneck is deciding the correct answers, not the technology, and that argument is worth having anyway.

### Do I have to redo it when the software updates?

Re-run it, yes — that is the point of having it. It takes minutes once it exists. Re-run the whole 200 whenever the underlying model changes, whenever you give the agent a new power, and whenever you change a policy such as your service-visit fee. Nothing else you can do gives you that much confidence for that little effort.

### What size of company does this make sense for?

If you have more than a few hundred systems in the field, your inbound volume is already past what one person can answer during install season. Below that, the honest answer is that a well-run answering service and a disciplined callback rule may serve you better until the fleet grows.

### Does this help with sales calls or only service?

Both, but they need separate test sets. A service call has a correct answer sitting in a system of record. A sales call is qualification: roof age, utility, ownership, whether there is shade, whether they are already talking to two other companies. Grade the sales side on whether it captured the fields your closer needs, not on what it said.

[CallSphere](https://callsphere.ai) builds AI voice and chat agents that answer business phone lines and web chat, capture the caller's details, and book appointments around the clock — which is exactly the stage-one and stage-two work described above. The part worth insisting on, from us or anyone else, is the folder of your own 200 calls and the grades before it goes live. An agent that answers quickly and confidently is not the achievement. An agent you have watched, step by step, on your own history is.

---

Source: https://callsphere.ai/blog/you-tried-a-chatbot-on-my-system-is-offline-calls-in-2024-build-the-20
