By Sagar Shankaran, Founder of CallSphere
How an industrial equipment OEM builds a parts-desk test set from its own closed orders, grades the agent on as-built revisions, and stops wrong-part shipments.
Key takeaways
The credit memo was on the controller's desk before the coffee finished brewing, and by then everybody in the building knew the story. A maintenance planner at a cheese plant outside Fresno had called the parts line about a machine your company built in 2013. Inside sales pulled the serial up in Epicor, found the drive-end seal kit on the as-shipped bill of materials, and put it on an overnight for $214 in freight. Wrong kit. That serial picked up a field change in 2017 that moved it to a larger shaft, and the seal went with it. The machine stayed down another day and a half, the customer got a goodwill credit, and your service manager spent forty minutes on the phone being polite.
Nobody was careless. The as-built history for that machine lives in three places: the original bill of materials in the ERP, the change order folder in SolidWorks PDM, and a field service report in a shared drive nobody has opened since the tech who wrote it left for a competitor. Every industrial equipment builder I have walked through has some version of this. It is also the exact job owners are now being told to hand to an agent — and this particular mistake is invisible until the return shows up.
Your aftermarket coordinator opens with roughly forty open lines and a voicemail box. The inquiries arrive three ways: the parts line, the shared parts inbox, and a photo texted to a regional service manager who forwards it with "customer needs this today." The typical message is not a part number. It is "we're leaking at the drive end, serial 4471, what do we need." Sometimes there is a nameplate photo, sometimes there is a nameplate photo of a different machine on the same line.
To answer it properly the coordinator opens the as-shipped bill of materials, checks the change order log for that serial, cross-references the gearmotor tag because the seal may belong to the gearbox vendor and not to you, checks the superseded-part table because the number on the 2013 drawing was replaced twice, and then calls the field tech who was last on that site. Fifteen to twenty-five minutes when it goes well. When it does not, it becomes a credit memo.
The workaround everyone pretends is fine: two people who have been there fifteen years know most of these answers cold. One of them is 61. Aftermarket parts carry the gross margin machine sales cannot, and the desk producing it is staffed like a cost center and documented like a hobby.
The 2024 version of this idea was a chat box bolted to your website that gave a fluent answer and no way to know whether it was right until the return authorization came back ninety days later. The 2025 version could search your documents, which helped, and still gave you no scoreboard.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent in your browser — 60 seconds, no signup.
What matured in 2026 is the boring half: measuring an agent against real cases and watching what it actually did, step by step, before you widen what it is allowed to touch. An agent test set is nothing more exotic than a folder of your own closed orders where you already know the correct answer, used to score a new system before it is ever allowed to speak to a customer. You run it unattended against those cases, you read the trail of what it opened and what it decided, and only then does it get to draft a reply — and later, maybe, send one.
flowchart TD
A["Pull 400 closed parts orders from the ERP"] --> B["Mark the part that stayed shipped, no return in 90 days"]
B --> C["Run the agent on each inquiry, unattended"]
C --> D{"Exact part number and correct as-built revision?"}
D -->|"Below the bar"| E["Read the step trail, fix the source record"]
E --> C
D -->|"At or above the bar"| F["Agent drafts, coordinator sends"]
F --> G["Re-score every Friday on that week's orders"]
G --> C
Take the last twelve to eighteen months of parts orders and pull three to five hundred lines where a human made a judgment and it held. "Correct" is not what somebody typed in the quote — it is the part that stayed shipped, with no return and no credit inside ninety days. That one rule turns your order history into an answer key without anyone writing an answer key.
Then load it with the ugly cases, because the easy ones prove nothing. Serials with field change orders. Bought-out components where the right answer is the gearmotor vendor's number and not yours. The two models you discontinued in 2016. Machines a customer's own maintenance team modified without telling you, which the field service reports will reveal. And twenty to thirty cases where the correct action was not an answer — it was "send me a nameplate photo" or "this needs the service manager."
Pull cases from the weeks that actually hurt: the July shutdown stretch and the two weeks before Christmas, when customers do their teardowns and your parts volume spikes while your desk is at half staff. An agent that scores well in April and falls apart on shutdown-week questions has not been tested.
Grade every case on five things, and grade them separately so you can see where it breaks:
Then read the step trail on a sample. You are looking for right answers reached for the wrong reason. An agent that ignored the change order log and happened to be correct on an unmodified machine is not accurate. It is lucky, and it will be unlucky on serial 4471.
Assumptions, stated plainly and illustrative: nine hundred parts orders a year through the desk, a current wrong-part rate of four percent, and a loaded internal labor cost of $70 an hour. One in six wrong shipments escalates into a downtime credit or a goodwill spare.
| Line | Today (4%) | After scoring (1.5%) |
|---|---|---|
| Wrong-part events per year | 36 | 14 |
| Outbound expedite + return freight, per event | $360 | $360 |
| Restocking and handling, per event | $95 | $95 |
| Coordinator 1.5 hrs + service manager 1 hr | $175 | $175 |
| Credit memo admin | $60 | $60 |
| Direct cost subtotal | $24,840 | $9,660 |
| Escalations at $1,200 each (1 in 6) | $7,200 | $2,400 |
| Total | $32,040 | $12,060 |
Roughly twenty thousand dollars a year on a desk that size, and it excludes the orders you never saw because the planner started calling a parts broker instead. Running an assistant across nine hundred text inquiries a year is cheap enough that freight dominates this table; frontier models are down roughly tenfold from 2025. The real build cost is your own people spending a few weeks cleaning the superseded-part table and getting the change order log into one place — work you needed anyway.
Still reading? Stop comparing — try CallSphere live.
CallSphere ships complete AI voice agents per industry — 14 tools for healthcare, 10 agents for real estate, 4 specialists for salons. See how it actually handles a call before you book a demo.
Do not let it commit a ship date on a long-lead item. Gearmotors, servo drives and variable frequency drives move on their own schedule, and a date promised by a machine is still a promise from you.
Do not let it answer on anything inside the safety circuit — light curtains, safety relays, interlock switches. Swapping a device changes the rating of that circuit and puts your risk assessment and your conformity file into play. That question goes to a named engineer, every time, no exceptions, and the agent's only correct move is to route it.
Do not let it decide warranty. Whether a bearing that failed at fourteen months is a warranty claim or an alignment problem at the customer's plant is a commercial judgment about a relationship, and your service manager gets paid to make it.
And do not point it at machines older than your records. If your as-built data starts in 2009 and the machine is from 1997, the honest answer is a photo and a phone call. The fifteen-year veteran does not get replaced by this. He gets interviewed by it — one hour a week on the cases the agent got wrong is the cheapest documentation project you will ever run.
Two to three hundred real closed orders beats a thousand invented ones. The mix matters more than the count: if every case is a current-model machine with clean records, the score is flattering and useless. Make a quarter of the set field-changed serials, discontinued models and bought-out components.
No, and trying to will kill the project. Run the scoring on the mess. The failures cluster, and they hand you the thirty or forty serial numbers whose records are actually costing you money. Clean those, re-score, move on.
Your aftermarket or parts manager owns it, and it is spreadsheet work, not engineering work. The one thing you should not do is hand it to applications engineering, who are already the bottleneck on quotes and will do it last.
It depends on how reversible the mistake is. Drafting a reply the coordinator sends is low risk at a modest score. Sending a part number that triggers an overnight shipment is a different bar, and most builders keep a human on the send button for two quarters, then remove it only for current-model, unmodified machines.
The same discipline applies to the parts line itself. When the coordinator is already on a call, the next planner with a down machine hears a voicemail greeting. CallSphere builds voice and chat agents that answer that line around the clock, take the serial number, the symptom and the site, and book the callback so the inquiry lands in your system instead of someone's memory. Score it the way you would score anything else that talks to your customers.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Community management agents fail on escalations and claims, not tone. Build a 400-thread test set from your own social inbox before anything goes live.
How a plant engineer builds the first list of machine builders in 2026, and the five specs an equipment OEM must publish as text to survive that pass.
Cattle producers can now measure a phone agent against real past calls, watch every step it took, and widen what it is allowed to do only once it earns it.
Sort six months of front-counter recordings into eight call types, write the right answer for each, then widen the agent's authority in four stages, not one.
How a design firm builds an RFI test set from closed jobs, scores citation and routing accuracy, and widens an agent's authority only when the numbers earn it.
Build a 300-call test set from your own cancellations and answering-service log, grade escalation at 100%, then widen the agent one visit type at a time.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI