By Sagar Shankaran, Founder of CallSphere
Routine colo tickets go to the fast cheap model and hard ones escalate. Your SLA severity matrix already drew that line. A worked monthly cost example inside.
Key takeaways
Three hundred and forty. That is a fair monthly ticket count for a colocation operator running around 900 cabinets across two sites — remote hands requests, access list changes, shipping and receiving, invoice questions, capacity enquiries, and the handful that make your operations manager stand up. Add roughly eleven thousand environmental and power alarm events on top, and you have the queue this post is about.
The question is not whether to put AI on that queue. Two years of arguing about that is over; US small-business adoption is at 66% and your customers are already using it on their side of the cage. The question is which of those 11,340 items gets the cheap fast model and which gets the expensive careful one — and, more importantly, who in your building decides where that line sits.
Cisco built model routing into the personal AI agent it is rolling out to roughly 90,000 employees, with an on-premises emphasis. The logic is the same at 90,000 seats and at nine: send the ordinary traffic to a fast, cheap model, and escalate only the genuinely hard cases to a strong one. In 2026 this stopped being a clever optimisation and became the default way anyone sane runs a queue.
Model routing means the fast cheap model handles the ordinary traffic while only genuinely hard cases go to the expensive one — and in colocation the line between the two is already written down, in the severity matrix in your master service agreement. You did not have to invent a taxonomy. Your lawyer and your first big customer negotiated one years ago.
A request is routine here when it is fully determined by records you already hold and it cannot take anything down. "What is my current draw on cabinet 22-11?" is routine — the metered rack power distribution unit knows, dcTrack knows, and reading it out changes nothing. "Add Marcus Deng to our access list for Thursday" is routine, provided your rule is that the authorised contact on file must be the one asking. "Send me last month's invoice with the cross-connect lines broken out" is routine. "Reboot the server in 12-04 via the console" is routine if the customer's standing authorisation covers it.
A case is hard when judgement, contract or risk enters. "Our circuit to Ashburn is flapping and we think your cross-connect is suspect" is hard, because the answer implicates your meet-me room and possibly a credit. "We need an emergency window during your December change freeze" is hard. "Can we add 12 kW to cage 4 by the end of the quarter?" is hard, because the honest answer depends on the branch circuits left on that remote power panel, the cooling in that row, and whether you want to sell that power to this customer at this price. And "we are claiming a service credit for the 14 minutes on the 9th" is hard by definition.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for IT support in your browser — 60 seconds, no signup.
flowchart TD
A["Ticket or alarm lands in the ConnectWise queue"] --> B{"Does it touch a live circuit, the fire system, or a credit?"}
B -->|No, routine| C["Fast model drafts the reply and updates dcTrack"]
B -->|Needs judgement| D["Strong model reads the MSA, change freeze and capacity plan"]
B -->|Always human| E["Page the chief engineer, no model action"]
C --> F["Remote hands tech executes, logs 15-minute increments"]
D --> G["Operations manager approves before anything is sent"]
Two people, and they are already on your org chart. The operations manager owns the routine side — which requests can be answered from records, which standing authorisations exist, what gets billed in fifteen-minute increments. The chief engineer owns the untouchable list — anything involving the fire suppression system, the paralleling gear, an automatic transfer switch, a loss of redundancy, or a customer on a single path.
The mechanism is a monthly sample audit, and it takes about an hour. Pull thirty routed items at random. For each one ask a single question: was it routed to the right place? Items that were routed cheap and should not have been are the only ones that matter; move that category up and write down why. Items routed expensive that were obviously routine are a cost story, not a risk story, and can wait. Keep the log — your auditor will eventually ask how automated decisions are governed, and Texas TRAIGA and California SB 53 have both been in force since 1 January 2026.
Illustrative rates. Frontier prices are down roughly tenfold since 2025, and a strong model still costs on the order of ten times what a fast one costs per item.
| Work type | Volume per month | Routes to | Cost each | Monthly |
| Environmental and power alarm classification | 11,000 | Fast | $0.004 | $44 |
| Access, billing, reading and shipping requests | 290 | Fast | $0.004 | $1 |
| Capacity, cross-connect faults, credits, freeze exceptions | 50 | Strong | $0.06 | $3 |
| Routed total | 11,340 | — | — | $48 |
| Everything on the strong model instead | 11,340 | Strong | $0.06 | $680 |
Look honestly at that table: at this volume the model bill is not the story. Six hundred dollars a month is a rounding error against one chiller compressor. The real result is what routing lets you do — run something on all eleven thousand alarms, which in 2024 you could not justify at all, and put the careful model on the fifty that are worth careful thought.
The labour number is where it moves. If classification collapses 11,000 raw alarm events into roughly 180 clusters that a human actually looks at, and each cluster previously cost a technician six minutes of reading, that is 1,082 hours a year of night-shift attention returned. At $62 loaded, call it $67,000 — not as a headcount cut, but as the reason your two critical facilities technicians can finally get through the preventive maintenance backlog.
Routing goes wrong in one direction financially: something starts escalating everything. A badly worded rule, a chatty customer thread, a loop where the strong model gets called on every reply in a chain. The Claude Enterprise governance update on 2 July 2026 added exactly the controls for this — a cost and usage dashboard, spend limits at the organisation and user level, alerts at 75% and 90% of budget, and model defaults you can hold people to. Set a monthly ceiling that is three times your expected bill and turn the alerts on. It costs nothing and it turns a surprise into an email.
Never cheap: anything quoting a price, anything referencing the service level agreement, anything a customer might later print out and bring to a credit negotiation, and anything about capacity you have not confirmed against the remote power panel. The cheap model is fast and fluent and will happily tell a customer there is room in a row that is full.
Still reading? Stop comparing — try CallSphere live.
See the IT support AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
Never routed at all: physical access decisions, fire and suppression events, anything that changes a breaker or a transfer switch position, and any communication during a live incident. During an outage your customers do not want a well-drafted message; they want the name of a person who is standing in the building. Give them that.
There is also a quieter limit. Routing works because the routine cases are genuinely determined by your records. If your records are wrong — if dcTrack has ports free that are physically occupied, if your access lists are eighteen months stale — then routing to the cheap model just makes you wrong faster. Fix the records first. That is not an AI project, it is a walk of the floor with a clipboard, and it is the single highest-return week of work available to most operators.
Yes, and most operators should. Alarms are internal, the outcome is measurable within a week, and nobody outside the building sees a mistake. Once the suppression and clustering numbers hold up for a quarter, move to the routine customer requests.
In the first ninety days, a human reviews everything before it sends — usually the operations coordinator, as part of the morning queue pass. After that, let the pure record-lookup categories send on their own and keep review on anything with a date, a price or a commitment in it.
Three artefacts: the written routing rules signed off by your operations manager and chief engineer, the monthly sample-audit log, and the spend and usage report. That is the same shape of evidence you already produce for change management, and it satisfies most of what state AI statutes and customer security questionnaires are actually asking.
Voice is its own decision. Speech-to-speech now answers in roughly two-tenths of a second and can look something up mid-sentence, which is fast enough to feel like a person. But apply the same line: a voice agent taking a cabinet access request is routine; a voice agent discussing an outage is not.
On that last point: CallSphere builds AI voice and chat agents that answer your main line and web chat 24/7, handle the routine end — who is calling, which cabinet, what they need, booking the escort or the site visit — and hand the hard ones straight to your on-call engineer with the details already captured. Same routing idea, applied to the phone that rings while both technicians are on the floor.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Sixty-one pages of addendum land three days before a DOT letting. How model routing gets an estimator the quantity changes that actually move the bid.
Model routing sends routine product requests to a cheap model and hard ones to a strong one. Here is who draws the line in a B2B software company, and how.
Cheap model for potholes, strong model for water quality, a pager for sewage in a basement. How a public works superintendent writes the triage table.
How water systems use a cheap-model first pass and a strong-model escalation so nobody drives out at 2 a.m. for a SCADA alarm that already cleared itself.
Sort courier calls before answering: cheap model for status and reschedules, strong model for damage claims and missed medical runs. Costs, rules and limits.
Colo cross-connect orders still arrive as faxed scans with handwritten port pairs. Here is what 2026 document reading does to the provisioning desk and the MMR.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI