By Sagar Shankaran, Founder of CallSphere
How a solar installer builds a test set from past service calls, watches the agent step by step, and widens its authority in stages without repeating 2024.
Key takeaways
Most solar companies that put a chatbot on the service line two years ago took it off within a quarter. It told a customer her system was fine when the gateway had been offline for nine days. It answered a question about the federal tax credit that your office manager had to walk back on the phone. It offered a Tuesday appointment in a ZIP code where you have no crew. Somebody printed the transcript, it got passed around the office, and that was the end of that conversation for two years.
Fair. But the reason it failed was not that the machine could not talk. It was that nobody could see what it did between the customer's question and its answer, and nobody had ever tested it against real calls with known right answers before it went live. That is the thing that actually matured in 2026. Watching an agent step by step and grading it against your own past cases became ordinary tooling, and it is now the gate that a rollout has to pass rather than an afterthought.
An agent evaluation for a solar installer is simply this: a folder of your own past customer calls with the correct answer written down for each one, run against the agent before it is allowed near a live line, and re-run every time anything changes. If a vendor cannot show you that folder and the grades, they are asking you to repeat 2024.
Pull a month of your call log out of RingCentral or Aircall and tag them by hand. In a residential solar and storage company with a few thousand systems in the field, they cluster hard:
Every one of these has a correct answer that already exists somewhere — in your job tracker, in the monitoring portal, in the signed agreement, in the utility's application status. And every one of them has a wrong answer that costs real money. That combination is exactly what makes them worth testing rather than guessing at.
Take the last 200 inbound calls or web chats. Not made-up examples — the real ones, including the angry one and the one where the customer was confused about which company installed the system. For each, write down what the agent should have said, what it should have checked first, and what it should have refused to answer.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for home services in your browser — 60 seconds, no signup.
That writing is the work, and it takes an office manager about six minutes a call. It is also, incidentally, the best service audit you will ever run, because halfway through you will discover that three of your staff answer the true-up question three different ways.
flowchart TD
A["Pull last 200 service calls from the ticket log"] --> B["Write the right answer for each one"]
B --> C["Run the agent on all 200, read every step"]
C --> D{"Did it ever guess at a bill, a guarantee or a credit?"}
D -->|Yes| E["Narrow what it may say, run the 200 again"]
E --> C
D -->|No| F["Live on answering and intake only"]
F --> G{"Two weeks clean on real calls?"}
G -->|No| E
G -->|Yes| H["Let it book the service visit"]
The final sentence the agent says is the least useful thing to grade. What you need to see is the sequence: did it identify the right site before it said anything about production, or did it match on a common last name? Did it actually read the monitoring status, or did it answer from the general shape of the question? When the customer said "my Powerwall," did it check whether that customer has storage at all, or did it play along?
That step-by-step record is what the 2026 tooling gives you and the 2024 chatbot did not. You can sit with your service manager and scroll through what the agent did on call 47, and she can point at the third step and say "that is where it went wrong — it should have asked for the service address first." Then you tighten that one rule and re-run all 200 in a few minutes rather than waiting for another customer to be the test.
Grade three things separately. Correct: it gave the answer your manager would have. Safe but incomplete: it did not know and handed off cleanly, which is a pass, not a failure. Unsafe: it stated something about money, coverage or equipment that it had no business stating. The unsafe count is the only one that has to be zero.
Stage one, read-only. It answers status questions and hands off everything else. Job status, install date, inspection scheduled, permission to operate received. No promises, no diagnosis, no dollars.
Stage two, intake and scheduling. It takes the service address, the system type, what the app is showing, whether there is a battery, and books a slot on a route day in the right region. It still refuses the money questions. This is where most of your saved hours are, because a full intake means your coordinator opens a ticket that is already complete instead of playing phone tag.
Stage three, narrow account actions, only if the earlier stages held for a month: sending the customer a copy of their agreement, confirming a lease transfer packet is on its way, resending a monitoring invitation. Each new power gets its own additions to the test set before it goes live. If you skip that step you are back to 2024, just with a better voice.
Assumptions, illustrative: 200 past calls, six minutes each to write the correct answer at a fully burdened $34 an hour, plus twelve hours of review time across two rounds of tightening. Live volume is 310 inbound calls a month.
Still reading? Stop comparing — try CallSphere live.
See the home services AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
| Item | Amount |
| Writing the answers: 20 hours | $680 |
| Two review rounds: 12 hours | $408 |
| One-time cost to build the test set | $1,088 |
| First run: unsafe answers | 9 of 200 (4.5%) |
| After tightening: unsafe answers | 0 of 200 |
Now price what you avoided. At a 4.5% unsafe rate on 310 calls a month, roughly 14 customers would have been told something wrong about a bill, a guarantee or their coverage. If one in five of those turns into a goodwill concession — a free service visit, a credit, a discounted cleaning to end the argument — at an average $900, that is about $2,520 a month you did not spend, against a one-time $1,088. And that is before the calls you miss entirely during install season.
Write this list before you write anything else, and check it in the test set explicitly. No statement about what a customer's utility bill will be, ever — rate structures and true-up cycles are where trust dies. No statement about tax credit or incentive eligibility; that is between the customer and their accountant, and the rules moved recently enough that even your closers get it wrong. No promise about how long a battery will run a house in an outage. No determination of whether something is covered under workmanship warranty versus the manufacturer's warranty, because that decision has your money attached to it.
And absolutely nothing that instructs a customer to touch equipment. Not the DC disconnect, not the rapid shutdown switch, not the breaker in the combiner. "Go flip the switch on the side of the inverter" is how a homeowner ends up in front of energised conductors, and no amount of testing makes that an acceptable thing for an automated voice to say. The safe version is: we will have a technician call you back, and here is when.
Keep a human on the angry call, too. A customer who has been waiting eleven weeks for permission to operate while making loan payments does not want an efficient answer; they want somebody with authority to apologise and commit to a date. Route those to a person by keyword and by tone, and check in the test set that the routing actually fires.
Two afternoons for the writing if your office manager blocks the time, plus a couple of hours per review round. The bottleneck is deciding the correct answers, not the technology, and that argument is worth having anyway.
Re-run it, yes — that is the point of having it. It takes minutes once it exists. Re-run the whole 200 whenever the underlying model changes, whenever you give the agent a new power, and whenever you change a policy such as your service-visit fee. Nothing else you can do gives you that much confidence for that little effort.
If you have more than a few hundred systems in the field, your inbound volume is already past what one person can answer during install season. Below that, the honest answer is that a well-run answering service and a disciplined callback rule may serve you better until the fleet grows.
Both, but they need separate test sets. A service call has a correct answer sitting in a system of record. A sales call is qualification: roof age, utility, ownership, whether there is shade, whether they are already talking to two other companies. Grade the sales side on whether it captured the fields your closer needs, not on what it said.
CallSphere builds AI voice and chat agents that answer business phone lines and web chat, capture the caller's details, and book appointments around the clock — which is exactly the stage-one and stage-two work described above. The part worth insisting on, from us or anyone else, is the folder of your own 200 calls and the grades before it goes live. An agent that answers quickly and confidently is not the achievement. An agent you have watched, step by step, on your own history is.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Supplement brands get adverse events reported by ticket, not by form. On-premises AI reads all 3,400 a month without that text ever leaving the building.
The unplanned failure that costs a solar installer most, the signals that precede it in Enphase and Powerwall data, and arithmetic on 1,800 warranty systems.
Vietnamese, Spanish and Mandarin-speaking dental front offices order the minimum and ask nothing. Live translation in 2026 changes the lunch-hour call.
The 200-call test set a treatment program should build from its own recordings, the five things to score, and how to widen an agent's authority safely.
How an industrial equipment OEM builds a parts-desk test set from its own closed orders, grades the agent on as-built revisions, and stops wrong-part shipments.
Community management agents fail on escalations and claims, not tone. Build a 400-thread test set from your own social inbox before anything goes live.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI