By Sagar Shankaran, Founder of CallSphere
Why per-seat AI pricing misfits utility work, when self-hosting an open model beats it for rate case data requests, and the staffing cost nobody budgets.
Key takeaways
Nine analysts in regulatory affairs. One general rate case. And north of a thousand written data requests arriving from commission staff, the office of consumer advocate and the industrial customer group between the filing and the close of the record — each due in ten business days, each needing a workpaper citation an auditor can follow, and a fair share needing a confidential-treatment or CEII stamp before they can be served.
Now price artificial intelligence against that. Every quote on your desk is per seat, per month. You have nine analysts, four rate people, two paralegals and a director. Fifteen seats. But the work is not fifteen people-shaped. It is two-thousand-documents-shaped, it lands in a nine-month lump, then goes quiet until the next filing three years out. Per-seat pricing bills you for the org chart. The rate case bills you for the pile.
Per-seat licensing is a fair deal when everybody uses the tool every day, all year. That describes a software company. It does not describe an electric utility.
Look at the three biggest bursts in your calendar. A general rate case: fifteen people, nine months, thousands of pages of discovery, then silence. A February ice storm: the call center goes from forty seats to three hundred inside twelve hours when you stand up the contracted overflow center, and it stays there for eleven days. A NERC audit cycle: two compliance analysts pulling six years of evidence across ninety days. Per-seat pricing makes you either buy for the peak and pay for it all year, or scramble to add seats in the middle of the event — which is precisely the week nobody has time to open a procurement.
An open-weight model is one the maker publishes in full, so a utility can run it on machines it owns and pay for hardware and electricity instead of paying a monthly fee for every employee who might open it. That is the whole difference, and in 2026 it stopped being a compromise.
The intervenor's third data request set arrives by email at 4:45 p.m. on a Friday, because it always does. Forty-one questions. Monday morning they get owners: rate design takes eleven, the depreciation study six, four to the vegetation management manager, three to the reliability engineer for SAIDI and SAIFI history, and two straight to counsel because they ask for substation drawings.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent in your browser — 60 seconds, no signup.
Then the real work starts, and it is not analysis. It is retrieval. Somebody spends two hours hunting for the answer the company gave to a nearly identical question in the last case, because the staff attorney recycles his own sets and the company must not contradict itself three years apart. Somebody else exports the continuing property records to tie a number back to what was filed on FERC Form 1. Somebody prints vegetation cycle spend by circuit and finds the 2023 figure in the workpaper does not match the operating budget, and now there is a reconciliation memo nobody planned for.
The workaround everyone pretends is fine has a name, and she has been in regulatory affairs nineteen years. She is the only person who knows the storm cost deferral workpapers from the last case live on a network drive under a folder named for a manager who retired. When she takes vacation mid-case, the response cycle slips two days.
Through 2024 and most of 2025, running your own model meant accepting a real quality gap. You could host something in your own data center, but the answers were worse than what analysts got from a browser tab, so they used the browser tab and the project died.
That gap closed this year. Moonshot AI published Kimi K3 in full — at 2.8 trillion, the largest open model in the world, and built so that only the slice of it a given question needs actually lights up, which is why it runs on hardware a utility can genuinely put in a rack. The whole open tier moved up with it, while the price of the licensed frontier services fell roughly tenfold from 2025. The choice is no longer good-and-rented versus poor-and-owned. It is: for this pile of work, which shape of bill do you want?
flowchart TD
A["Intervenor data request set lands 4:45pm Friday"] --> B["Model drafts each response from prior case answers and workpapers"]
B --> C["Analyst checks every citation against the source workpaper"]
C --> D{"Confidential or CEII treatment needed?"}
D -->|Yes| E["Counsel stamps and redacts before service"]
D -->|No| F["Regulatory affairs serves the response"]
E --> F
F --> G["Answer filed into the case library, indexed by question"]
G --> B
The analyst opens the case library, which now holds every data request response the company has served in the last three cases, the workpapers behind them, the FERC Form 1 filings and the cost-of-service study. She pastes in the new question — "Provide by circuit the vegetation management expenditures for 2022 through 2025 and reconcile to Exhibit RJM-4" — and gets a draft in ninety seconds, with the prior case's answer to the near-twin question flagged and every figure carrying the workpaper tab it came from.
She still does the reconciliation herself, because the 2023 mismatch is real and no model is going to know the difference is a storm-restoration tree crew miscoded to O&M. But she starts from a draft with the citations assembled rather than from a blank page and a network drive. The response goes out on day six instead of day nine, which matters because the next set is already queued behind it.
Take an illustrative mid-size investor-owned utility. Not your numbers — put your own in the same frame.
| Assumption | Value |
| Regulatory, rates, compliance and legal seats, all year | 15 |
| Call center seats needed only in storm season | 60 for 6 weeks |
| Illustrative business-tier seat price | $220 per month |
| Per-seat, buying for the peak all year (75 seats) | $198,000 per year |
| Per-seat, 15 all year plus 60 for 6 weeks | $58,140 per year |
| Self-hosted: two server nodes, over 3 years | $41,000 per year |
| Power and cooling, at your own industrial rate | $6,800 per year |
| People: 0.5 systems engineer, 0.25 records | $118,000 per year |
| Self-hosted total | $165,800 per year |
Read that honestly. Against buying 75 seats all year, running your own wins by roughly $32,000. Against buying 15 seats and adding the storm surge only when a storm actually arrives, per-seat wins by more than $100,000. Self-hosting gets strong when the seat count climbs past about seventy, when the same machines also serve the compliance team searching six years of audit evidence, or when a contract or a regulator says the data must not leave the building — which for parts of a utility, it does.
Still reading? Stop comparing — try CallSphere live.
CallSphere ships complete AI voice agents per industry — 14 tools for healthcare, 10 agents for real estate, 4 specialists for salons. See how it actually handles a call before you book a demo.
Running your own is not free of people, and the people are the part that sinks projects. You need someone who owns the machines — patching, capacity, keeping the thing up when the rate case team is filing at 11 p.m. on a deadline night. Half a full-time person, and it cannot be the same overloaded person who owns the outage management system.
You also need someone who owns the library. Every question the model answers well is a question whose source document was findable, current and correctly labelled. That is a records job, not a computer job, and in most utilities it is nobody's job today. Give it a quarter of somebody in regulatory affairs and the whole thing works. Give it to nobody and you have built an expensive way to get confident wrong answers about vegetation spending in 2023.
If your regulatory calendar is quiet — no rate case pending, no storm cost recovery filing, no depreciation study — buy seats for the handful of people who need them and skip the rack. Self-hosting is justified by volume and by data rules, not by principle.
And nothing here signs anything. A data request response is a sworn statement in an evidentiary record. The analyst who signs it owns every number in it, and if a figure is wrong, the fact that a draft produced it is not a defence in front of an administrative law judge. Keep counsel between the draft and the service list, and never let anything with a CEII stamp move without the person named on the protective order agreeing to it.
No. Two well-specified servers in a rack you already own will carry a regulatory team. What you do need is a real owner for those machines and a clear answer to where they sit relative to your corporate network and your control-centre network, because those are not the same place and your compliance manager will ask.
Then leave it alone for the office work. The argument for running your own is narrower than the vendors on either side make it sound: it is for high-volume, bursty, document-heavy work and for information you have contractual or regulatory reasons to keep on your own premises. If neither applies to you this year, the answer is genuinely to do nothing.
Take the last completed rate case. Load every data request and response into one searchable library with the workpapers attached, and measure how long it takes an analyst to answer ten questions from the current set with it versus without it. That measurement is worth more than any vendor demonstration, and you can start it Monday with the case files you already have.
In practice it does not, because the commission does not send fewer questions. What it changes is the response cycle and the dependence on the one person who knows where everything is filed — a continuity problem worth solving before she retires.
One footnote from the other side of the building: the week a new rate takes effect on bills, the call centre gets hit harder than the rate case ever hit regulatory affairs. CallSphere builds AI voice and chat agents that answer business phone lines and web chat, book appointments and capture leads around the clock — useful for the bill-explanation and appointment calls that pile up after a rate change, while your people stay on the accounts that need a human.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
When per-seat AI licensing stops paying for a benefits agency, why the claims files decide it, and the three-year arithmetic on owning a machine instead.
The monthly IEEE 1366 reliability close takes 64 hours across three people. What goal-driven agents change, the arithmetic, and what stays with the engineer.
Open models closed the gap in 2026. A worked buy-versus-rent comparison for a 58-employee credit union, including the staffing cost nobody budgets for.
400 estoppels and 1,900 inspection photos a month change the math on AI licensing for community association managers. A worked breakeven, plus the limits.
When per-seat AI pricing stops making sense for a host agency or tour operator, what running an open model actually costs in people, and the data case for it.
Kimi K3 and the 2026 open tier let a building materials dealer host instead of license per seat. Worked cost comparison for 4,000 January cost changes.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI