By Sagar Shankaran, Founder of CallSphere
Molding quotes fail on cycle time and scrap. Build a test set from closed jobs with monitored actuals, grade the agent on four marks, then widen what it may do.
Key takeaways
You tried this. Somewhere in 2024 a vendor put a quoting assistant in front of your estimator, it looked at a print, and it came back with a 19-second cycle on a part that ran 26 on the floor. Nobody could tell you why it said 19. Your estimator went back to his spreadsheet and the whole thing died in six weeks. That was the right call at the time.
What changed in 2026 is not that the agents got magically better at reading prints. It is that the tooling for watching and grading an agent grew up. You can now run an agent against a pile of jobs you already know the answer to, see every step it took to reach its number, and only widen what it is allowed to do once it earns it. The gate is measurement, not faith — and measurement is something a molding shop is unusually well equipped to do, because you already have the answers sitting in your production monitoring history.
An injection molding quote is not one estimate, it is a chain of them, and the errors compound. Cycle time is first and worst: cooling dominates, wall thickness drives cooling, and a half-millimeter of nominal wall on a PC part is the difference between 24 and 31 seconds. Then cavitation — quoting a 4-cavity tool the customer eventually buys as a 2-cavity doubles your machine hours per thousand. Then the scrap and startup allowance, which everyone sets at 2% because that is what the template says, on a part that historically runs 5.5% because of gate blush on the B-side. Then secondary operations: the pad print that adds 11 seconds of labor per part and never made it onto the router.
Miss cycle by 20% on a 1.2-million-piece annual program and you have given away roughly a fifth of your machine revenue on that part for four years. That is the risk anyone is really talking about when they say they do not trust an AI quote.
Through the first half of 2026, the practical part of the agent story stopped being the model and started being the instruments around it. You can record every job an agent handles, replay exactly what it looked at and in what order, score its output against a known-correct answer, and hold it at a narrow scope until the score holds. Enterprises doing this at scale are not turning agents loose; they run them shadow-mode against real history first and widen authority in steps.
A test set, in this trade, means a batch of your own closed molding jobs where you already know the true cycle time from the machine monitor, the true scrap rate from the job history, and whether you won or lost the quote — used to grade an agent before it is ever allowed to answer a customer.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent in your browser — 60 seconds, no signup.
flowchart TD
A["Pull 200 closed jobs from DELMIAworks or Plex"] --> B["Attach the real cycle from the machine monitor, not the router"]
B --> C["Agent re-quotes each job with the customer name hidden"]
C --> D{"Cycle within 10% and scrap within 1 point?"}
D -->|No| E["Estimator replays the agent's steps and fixes the rule it used"]
E --> C
D -->|Passes 9 of 10| F["Agent drafts quotes under 250,000 pieces a year only"]
F --> G["Quality manager re-scores the same 200 jobs every 90 days"]
Most shops already have everything they need and have never assembled it in one place. If you run DELMIAworks, Plex, Global Shop, or E2, the closed job history has the estimate. If you run Mattec or any machine monitoring at all, the presses have the truth. The job is to sit them next to each other.
Pull 200 jobs closed in the last three years. For each one, capture: the quoted cycle and the actual average cycle from the monitor, not the standard on the router — those two drift apart the moment a process tech dials in a real setup. Capture the quoted scrap allowance and the actual scrap from the job. Capture the quoted price per piece and whether you won it. Capture the resin, shot weight, cavitation, tonnage, and whether the tool was yours, the customer's, or offshore-built.
Weight the set deliberately. Half should be your bread and butter — the parts you quote in your sleep. A quarter should be the ugly ones: thick sections, hot runner with valve gates, insert molding with a brass insert an operator loads by hand, anything with a texture spec. And a quarter should be jobs you lost, with the price you lost at, because an agent that quotes accurately but always high is a different failure than one that quotes low.
Score it on four things and do not let anyone add a fifth until those four are green:
| Grade | Pass mark | Why this one |
| Cycle time | Within 10% of the monitored actual on 9 jobs out of 10 | Cycle drives machine hours, which drives everything |
| Scrap allowance | Within 1 percentage point of actual | Catches parts with known cosmetic problems |
| Secondary operations | Every operation on the router is named, none invented | Forgotten pad print and heat stake eat the margin quietly |
| Price posture | No more than 5% below the price you actually won at | An agent that buys work is worse than no agent |
The part that is new and worth the effort: when it fails a job, you open that job and read what it did. Usually it is one bad rule — it used nominal wall from the print instead of the thickest section, or it applied your standard 22-second cooling rule to a filled nylon that behaves nothing like unfilled PP. You fix the rule, re-run all 200, and see the score move. That loop used to be guesswork.
Assume you quote 34 programs a year and win 8. Assume the average awarded program is 600,000 pieces a year at 22 seconds on a 300-ton press, and you cost machine time at $62 per hour. A 20% cycle miss on one awarded program is worth roughly 3,700 hours of machine time you did not price over a four-year life at 600,000 pieces a year — but let us keep it to a single year to stay honest.
| Annual volume, one program | 600,000 pieces |
| Quoted cycle, 2-cavity tool | 22 seconds, 300 pieces per hour |
| Actual cycle | 26.4 seconds, 250 pieces per hour |
| Machine hours quoted vs actual | 2,000 vs 2,400 hours |
| Unpriced machine hours at $62 | $24,800 per year |
| Over a four-year program | $99,200 |
| Cost of the test set | Roughly 20 hours of the estimator's and quality manager's time to assemble |
One caught cycle error on one program pays for the exercise several times over — and the same 200 jobs also tell you how often your human estimator misses, which is the number nobody has ever measured in most shops.
Still reading? Stop comparing — try CallSphere live.
CallSphere ships complete AI voice agents per industry — 14 tools for healthcare, 10 agents for real estate, 4 specialists for salons. See how it actually handles a call before you book a demo.
Keep a human on anything the test set does not cover. A resin you have never run in production. A tool coming from a mold shop you have not worked with, where the real risk is steel condition and water line layout, not geometry. Anything going into a validated medical process where the quote implies IQ, OQ and PQ runs and a customer-witnessed capability study. Anything where the customer is supplying the tool and you have not seen it — the quote is really a bet on someone else's steel.
Keep a human on the strategic price too. Sometimes you quote a part at a number that makes no sense on the spreadsheet because it holds a family of six other part numbers on your floor. No test set contains that reasoning, and it should not.
And keep a human on the conversation. When the buyer calls to say the price is 9% high, the answer is usually a change to cavitation, a different runner, or a volume commitment — a negotiation, not a recalculation.
Ninety with trustworthy monitored cycle data beats 400 where the actuals are the router standard copied forward. Start with what is clean, and add jobs as you close them. The one thing you cannot compromise on is that the actual cycle comes from the press, not from the estimate.
Then build the test set from those twelve and be explicit that the agent is only graded on that tonnage range. Scope the authority to what you measured. That is exactly the discipline this approach is for.
Those are the most valuable jobs in the set. Record both the original and the improved cycle. If the agent consistently quotes the improved cycle you have an optimist; if it quotes the original you have a conservative. Either is fine as long as you know which one you have.
Not on the first pass. Start with drafting quotes that your estimator signs, on programs under a volume threshold you set. Widen it after two quarters of scores that hold. The whole point is that authority follows evidence.
Faster quoting only helps if you can also handle the calls that come with it — the buyer chasing a number, the customer asking whether you can hold the price if the volume doubles, the follow-up nobody made because everyone was on the floor. CallSphere builds AI voice and chat agents that answer the phone line and the web chat, capture who called and what program they were asking about, and book the follow-up on your estimator's calendar. It does not price your parts — your estimator and your test set do that.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Sixty-one pages of addendum land three days before a DOT letting. How model routing gets an estimator the quantity changes that actually move the bid.
How an industrial equipment OEM builds a parts-desk test set from its own closed orders, grades the agent on as-built revisions, and stops wrong-part shipments.
Community management agents fail on escalations and claims, not tone. Build a 400-thread test set from your own social inbox before anything goes live.
Cattle producers can now measure a phone agent against real past calls, watch every step it took, and widen what it is allowed to do only once it earns it.
Sort six months of front-counter recordings into eight call types, write the right answer for each, then widen the agent's authority in four stages, not one.
How a design firm builds an RFI test set from closed jobs, scores citation and routing accuracy, and widens an agent's authority only when the numbers earn it.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI