By Sagar Shankaran, Founder of CallSphere
Routing sends routine SOC alerts to a cheap model and hard ones to a strong model and a named analyst. Where the line sits, who moves it, the monthly numbers.
Key takeaways
You tried this in 2024. A vendor sold you automated triage, it auto-closed a batch of alerts overnight, and three weeks later a Tier 3 analyst found one of them was the first sign of a service account being abused at a client. You turned it off, told the team never again, and went back to humans reading every alert. Fair enough.
What changed since is not "the models got smarter". It is that routing became normal practice: instead of one model handling everything, routine work goes to a fast cheap model and only genuinely hard cases escalate. Cisco built exactly this into the personal AI agent it is rolling out to roughly 90,000 employees, to balance cost against capability. It matters in a SOC because the 2024 failure was never a model problem — nobody drew the line between routine and hard, wrote it down, and put a name against it.
Pull a month of triage tickets and sort them by detection rule. In most managed detection practices the shape is the same: the large majority of what reaches a human is a handful of recurring patterns. Impossible-travel sign-ins from a client whose sales team uses a consumer VPN on hotel wifi. The EDR agent quarantining the network scanner the client's own audit team runs on a schedule. A nightly backup job tripping a credential-access rule. A shared mailbox rule that fires whenever the accounts payable clerk sets an out-of-office.
Each takes an analyst four to eight minutes: open the ticket, pivot into the console, check the client's tuning notes in the runbook, confirm it matches a decision made three months ago, write two sentences, close. Not hard, not interesting, and 85% of a Tier 1 shift — which is why turnover runs as it does and why you re-explain the same client quirk every eight months.
Model routing in a SOC means the routine, repeatedly-seen alerts get handled by a cheap fast model under a human's review, while anything novel, multi-signal or high-consequence goes straight to the strong model and then to a named analyst.
An alert is routine when three things are true at once: it matches a detection rule you have tuned for this client before, it involves one signal rather than a chain, and the affected account and device are ordinary — standard user, managed laptop, no privileged group, nothing in a regulated enclave. The work is then mechanical: gather the context, compare against the documented decision, propose the closure.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for IT support in your browser — 60 seconds, no signup.
An alert is hard the moment any one of these appears. It touches a domain controller, a backup server, a hypervisor host, or anything inside a CUI enclave at a defense client. It involves a privileged or service account. It is two or more signals in sequence — a sign-in followed by a mailbox rule, an execution followed by outbound traffic to somewhere new. It fires on a rule nobody has tuned for this client. Or the client is one whose reporting obligations start the moment you confirm an incident.
None of those conditions is about the model's confidence score. Confidence is the wrong dial: a cheap model is perfectly confident about a novel attack it has never seen described. Route on the properties of the alert and the client, not on how sure the machine says it is.
flowchart TD
A["Alert lands in the triage queue"] --> B["Router checks rule, account type, asset, client tier"]
B --> C["Lane 1: seen before, one signal, standard user"]
B --> D["Lane 2: privileged account, server, or new rule"]
B --> E["Lane 3: multi-signal chain or regulated client"]
C --> F["Cheap model drafts closure, analyst reviews in 90 seconds"]
D --> G["Strong model builds the timeline, Tier 2 decides"]
E --> H["Tier 3 and the incident lead, client called"]
F --> I["Weekly tuning meeting re-reads 25 random closures"]
G --> I
This is what failed in 2024, and it is a management question, not a technical one. The line between routine and hard is a written escalation policy owned by two named people: your detection engineering lead, who knows what each rule fires on, and your SOC manager, who owns the service levels in the client contracts. It gets versioned like any other document, with a date and an author.
It gets reviewed on a fixed cadence — a Thursday tuning meeting works — with the service delivery manager in the room, because moving a rule from hard to routine changes what a client's monthly report looks like. Every routing change gets recorded with a reason. When a client asks in their quarterly review why an alert was closed without a human writing the summary, you want an answer with a date on it.
What must never happen is a vendor's default settings quietly deciding your escalation policy. If the tool ships with its own routing logic, either you can see and edit the conditions or you treat everything it handles as unreviewed. The clause in your own contracts — that a qualified analyst reviews detections — is a promise you made, and it does not transfer to a supplier.
Illustrative assumptions: 62 client tenants, 4,800 alerts a month reaching human triage, 88% of them routine. Analyst cost $68 an hour fully loaded. Today an average triage takes 6 minutes. With a drafted summary in front of them, a routine review takes 90 seconds; hard alerts get longer, not shorter, because the timeline is already built and the time goes on judgment.
| Today | With routing | |
|---|---|---|
| Routine alerts (4,224) | 6 min each = 422 hrs | 1.5 min each = 106 hrs |
| Hard alerts (576) | 6 min each = 58 hrs | 11 min each = 106 hrs |
| Total analyst hours per month | 480 hrs | 212 hrs |
| Analyst cost per month | $32,640 | $14,416 |
| Model cost (illustrative: 2¢ routine, 35¢ hard) | $0 | $286 |
| Monthly difference | $17,938 |
Two honest notes. First, the saving is not a layoff — 268 hours a month is roughly 1.6 analysts' worth of time, and in most firms it gets spent on the threat hunting and detection tuning you have promised clients since the last renewal, which is what justifies your retainer. Second, the cheap lane costs pennies because model prices fell roughly tenfold from 2025. Running a model over every alert now costs less than the SOC's monthly coffee order, which was not true when you tried this two years ago.
The 2024 failure mode is still available to you. Routing narrows it, because the dangerous categories never enter the cheap lane, but a novel attack that happens to look like a tuned-out pattern can still get a machine-drafted closure and a distracted human clicking approve at 4 a.m.
Still reading? Stop comparing — try CallSphere live.
See the IT support AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
So build the audit in from day one: every week someone senior re-reads 25 randomly chosen routine closures cold, without seeing the draft summary first. Track the disagreement rate. Above a couple of percent, the routing conditions are too loose and something moves back to the hard lane. That review is thirty minutes a week, and it is the reason you can tell a client with a straight face that a person is accountable for every closure.
Keep humans on three other things. Declaring an incident and starting a client's notification clock is yours and stays yours. The client phone call — telling a CFO their controller's mailbox has rules forwarding invoices — is a relationship task, not a summary task. And tuning stays with your detection engineer; a model can flag that a rule is noisy, but somebody who has met that client decides whether the noise is a bad rule or a bad habit worth raising.
Do not route the queue. Pick the single noisiest detection rule across your book — for most firms an identity rule — and one client. Have the cheap model draft closure summaries for that rule and that client only, with a human approving every one. Record how long review takes and how often the analyst disagrees. Two weeks of that gives your SOC manager a real disagreement rate and the evidence to write the first escalation policy with numbers in it rather than a vendor's promise. Then add the second rule.
Not if a human still approves every closure, which is the design here — the model drafts, a person decides. What would break it is auto-closure with sampled review, which some firms do sell at a lower price with the contract written accordingly. Read your own agreement first, and if you move to sampling later, price it as a different service tier rather than quietly redefining the one clients already bought.
Ask the question your CMMC-scoped clients will ask you: which service handles the alert content, where does it run, is it covered by your existing agreements, and does it sit inside the boundary described in your own system security plan. For clients with controlled unclassified information, the safe answer is to keep their alerts out of the routing entirely, or to use an installation your assessor has already accepted.
Some will, which is why the weekly cold re-read of random closures exists and why the disagreement rate is a metric you watch by analyst, not just in total. If one person never disagrees with a draft and everyone else disagrees three percent of the time, you have a coaching conversation, not a tooling problem.
Realistically it lets the same triage team absorb the next eight to ten tenants, provided their alert mix resembles your existing book. It does not scale the hard lane, and the hard lane is where your senior people live. If you are winning larger or more regulated clients, hiring Tier 3 is still the constraint.
Triage is not the only queue with a routine majority. The other is the phone: the client asking whether last night's alert email needs action, the prospect wanting a scoping call, the questionnaire chaser. CallSphere builds AI voice and chat agents that answer the line and website chat 24/7, handle routine questions, book the call and capture the lead into your systems — while anything that sounds like an actual incident goes straight to your on-call analyst instead of a voicemail box.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Sixty-one pages of addendum land three days before a DOT letting. How model routing gets an estimator the quantity changes that actually move the bid.
Model routing sends routine product requests to a cheap model and hard ones to a strong one. Here is who draws the line in a B2B software company, and how.
Cheap model for potholes, strong model for water quality, a pager for sewage in a basement. How a public works superintendent writes the triage table.
How water systems use a cheap-model first pass and a strong-model escalation so nobody drives out at 2 a.m. for a SCADA alarm that already cleared itself.
Sort courier calls before answering: cheap model for status and reschedules, strong model for damage claims and missed medical runs. Costs, rules and limits.
Six tests that make a title order routine, the escalation list that never bends, who owns the rule, and a costed 90-file month showing where the money sits.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI