By Sagar Shankaran, Founder of CallSphere
How water systems use a cheap-model first pass and a strong-model escalation so nobody drives out at 2 a.m. for a SCADA alarm that already cleared itself.
Key takeaways
The callout phone goes off at 2:14 a.m. It is the alarm dialer reading a SCADA point aloud in that flat synthetic voice: high discharge pressure, Route 9 booster. The on-call operator acknowledges from bed, drives nineteen minutes, badges into the station, and finds a pump that already restarted itself and a discharge reading four PSI over setpoint because the pressure-reducing valve on the east side hunted for ten minutes and settled. He resets nothing, drives home, and puts two hours of callout pay on the timesheet, which is fair, because he was awake and on the road.
Three weeks later the same phone rings at 3:40 a.m. with low free chlorine residual at the north entry point. Same voice, same three-word alarm text, an entirely different night — a Tier 1 public notice and a boil-water advisory if it is real, and the one where the twenty minutes between the page and the operator's boots on the floor actually matter.
This is a sorting problem, not a technology problem. The dialer has no idea which of those two nights it is looking at, so the callout list is set conservatively, everything wakes somebody, and operators learn to answer the phone with their eyes closed.
The setup is close to identical everywhere. SCADA runs on VTScada, Ignition, Survalent or an older Wonderware build. Alarms leave through a dialer or a cellular unit — Mission, High Tide, WIN-911 — and land on a rotating on-call phone. Callouts pay time-and-a-half with a two-hour minimum whether the operator was there ninety seconds or ninety minutes.
The workaround everybody pretends is fine is the mental filter. The Grade III operator who has been there eleven years knows the Route 9 high-pressure alarm at low demand is almost always the PRV hunting. He drives out anyway, because the one time he does not will be the night the check valve is stuck and the tank overflows down the side of the standpipe. That knowledge lives in one man's head and leaves with him when he retires.
Model routing means the fast, cheap model handles every alarm matching a pattern it has seen resolve safely, and only the ones it cannot confidently place go up to the expensive model that reasons through tank levels, residuals and pressures together. It is a sorting step in front of the callout list, not a decision-maker on top of it. Cisco built exactly this into the personal AI agent it is rolling out to roughly 90,000 employees — routine to the fast model, hard ones escalated — for the reason a utility would: running the heavyweight on everything costs money and buys nothing on the easy ninety percent.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent in your browser — 60 seconds, no signup.
Alarm rationalization is not new — dead-band tuning, shelving and priority tiers predate anybody saying "AI." What is new in 2026 is that the second stage is affordable and good enough to trust with a judgement call. You can now hand the strong model the last thirty minutes of tank level, entry point residual, finished water flow and pressure at three points and let it reason about whether those four things tell a coherent story.
Where it runs changed too. A first-pass sorter on a small server in the control room keeps working when the fiber to the plant is cut, and only ambiguous cases go out to the big model — a design a superintendent can explain to a board and to the state inspector.
flowchart TD
A["SCADA alarm at 2:14 a.m."] --> B["Fast model reads last 30 minutes of trends"]
B --> C{"Does this match a pattern that cleared itself before?"}
C -->|Yes, pump restarted, pressure settled| D["Log it, no callout, note on the morning report"]
C -->|No, or residual and turbidity involved| E["Strong model reviews tank level, residual, flow, pressure together"]
E --> F{"Could this be a public health event?"}
F -->|Yes| G["Call the on-call operator and page the ORC"]
F -->|Not now, but abnormal| H["Hold for the 6 a.m. shift, open a work order in Cityworks"]
G --> I["Operator confirms on site, sampling kit in the truck"]
11:52 p.m. Low suction pressure at the Mill Street lift station, cleared in forty seconds — it happens whenever the grinder pump at the apartment complex dumps, and the fast model has seen it 300 times. No callout. It becomes a line on the morning report the lead operator reads at 6:10 a.m.
1:31 a.m. Standpipe level dropping 1.4 feet per hour with no pumps calling. No clean match, so the case goes up. The strong model compares the trend to the last fourteen Tuesdays, sees district flow up 380 gallons per minute in the southwest zone, and writes one paragraph: looks like a main break in the southwest pressure zone, not an instrument fault, recommend callout. The operator gets a call that says what is probably happening and why, instead of three words and a point name.
4:05 a.m. Entry point residual drifting from 0.9 to 0.6 mg/L. Anything touching disinfection escalates automatically no matter what either model thinks. Phone rings, ORC is paged, and the operator is pulling a sample at the first site on the state-approved sample siting plan before sunrise.
Assume a 6,000-connection system, three operators on a weekly rotation, 41 after-hours callouts in a quarter, and that 22 of them turn out to be self-clearing conditions already visible on the trend before the truck left the driveway. Loaded overtime is $63 an hour with a two-hour minimum and the round trip averages 26 miles. These are illustration numbers — pull your own from the callout log and the timesheets.
| Line | Today | With routing |
| After-hours callouts per quarter | 41 | 21 |
| Callouts later judged self-clearing | 22 | 2 |
| Overtime paid (2 hr min at $63/hr) | $5,166 | $2,646 |
| Mileage (26 mi round trip) | $779 | $399 |
| Quarterly total | $5,945 | $3,045 |
That is roughly $11,600 a year, which does not buy a new pump. The number that matters more is not on the table: the operator who did not lose three hours of sleep is sharp on Wednesday, and the callout that does come carries a note explaining itself, so he brings the right truck and the right sampling kit.
Somebody has to write down which alarms may be sorted and which always ring the phone, and that somebody is the operator in responsible charge, not the software company and not the IT contractor. In practice it is a two-column list attached to the standard operating procedures: alarms eligible for suppression, and alarms that escalate unconditionally.
Still reading? Stop comparing — try CallSphere live.
CallSphere ships complete AI voice agents per industry — 14 tools for healthcare, 10 agents for real estate, 4 specialists for salons. See how it actually handles a call before you book a demo.
The unconditional column usually reads: any disinfection residual out of range, any turbidity excursion at a filter, any tank below the low-low setpoint, any pressure below 20 PSI in distribution, any intrusion alarm, any generator failure to start during an outage. The eligible column is the noise — pump fail-to-start that restarts, communication dropouts under five minutes, high-level alarms during a fill cycle, door alarms in business hours. Keep the log either way: when the state inspector asks how you handle alarms after hours, the answer needs to be a document with the ORC's name on it and a record of every alarm, including the ones nobody was called for.
The sorter has no idea what happened outside the fence. It does not know a contractor hit a 6-inch main on Cypress at 9 p.m., that the fire department drafted from a hydrant on the east side, or that the school is filling a pool. Weather, digs and fires are what make a normal-looking trend abnormal, and the operator knows all three.
It also cannot judge a bad instrument. A residual analyzer with a fouled cell reads low and looks exactly like a real loss of residual; only somebody in front of it with a bench colorimeter can tell. That is why residual escalates unconditionally — the model is not being asked to decide, only to write a better message than three words.
And it should never move a valve, change a setpoint or start a pump. Read-only, off a mirrored copy of the historian, is the only setup worth defending in front of your instrumentation tech and your insurance carrier.
Not for the alarm itself. Sanitary surveys look at whether you have written procedures, whether you follow them, and whether monitoring and reporting were complete. A documented alarm response procedure signed by the ORC, with a log of every alarm and its disposition, is stronger evidence of control than a callout list that wakes a man for a door contact.
No. Every one of these systems can log to a historian or a database, and that copy is what gets read — nothing touches the control network in the write direction. If an integrator tells you the only path is a full upgrade, get a second quote.
It will, eventually. Build the review into the Monday morning meeting: the lead operator reads the past week's suppressed list, and anything that looks wrong moves to the unconditional column permanently. Ask any vendor for a monthly spending ceiling and an alert when you approach it — the enterprise tools added exactly those controls in July 2026.
Export twelve months of alarm history out of SCADA into a spreadsheet and write, next to each line, what the operator actually found. That file — alarm, time, disposition — is the whole project. Most superintendents who do it find that six point names generate more than half the after-hours calls, and that fixing two instruments and one PRV setpoint captures much of the benefit before anybody buys anything.
One last thing worth naming: every main break and boil-water notice generates a wall of calls to the office line, at the exact hour the crew is in the street and nobody is at a desk. CallSphere builds AI voice and chat agents that answer the utility's public line around the clock — taking the address, the nature of the complaint and a callback number, telling callers what the district has already posted, and passing anything urgent to the on-call operator. It does not read your SCADA screen and it does not decide what is an emergency. It keeps the customer side of a 2 a.m. event off the person standing in the hole.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Sixty-one pages of addendum land three days before a DOT letting. How model routing gets an estimator the quantity changes that actually move the bid.
Model routing sends routine product requests to a cheap model and hard ones to a strong one. Here is who draws the line in a B2B software company, and how.
Cheap model for potholes, strong model for water quality, a pager for sewage in a basement. How a public works superintendent writes the triage table.
Sort courier calls before answering: cheap model for status and reschedules, strong model for damage claims and missed medical runs. Costs, rules and limits.
A monthly drone flight over the standpipe, lagoon berms and booster station catches the deficiencies a sanitary survey would otherwise find three years late.
Six tests that make a title order routine, the escalation list that never bends, who owns the rule, and a costed 90-file month showing where the money sits.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI