By Sagar Shankaran, Founder of CallSphere
When per-seat AI licensing stops paying for a benefits agency, why the claims files decide it, and the three-year arithmetic on owning a machine instead.
Key takeaways
How many of your people need AI every single day, and how many pages need it once a year? Those are two completely different purchases, and the per-seat price list is designed so you only ever think about the first one.
A 25-person benefits agency has maybe nine account managers, four account executives, two analysts, a compliance manager, an enrollment specialist and a handful of producers. At $30 a seat a month, giving everyone a licence costs $9,000 a year, which is nothing. The problem is not the nine account managers writing better emails. The problem is the 9,000 pages that arrive between mid-September and the first of December, and the fact that most of those pages contain somebody's claims history.
Per-seat pricing works beautifully when the work is one person, one task, one moment: an account executive drafting a renewal summary for a client meeting, a compliance manager rewriting a summary plan description paragraph, a producer turning a discovery call into a proposal outline. Nobody should build anything to do that. Buy the seats.
Per-seat pricing falls apart the moment the unit of work stops being a person and starts being a document count. Reading every summary of benefits and coverage in the book, side by side with last year's, is not nine people doing nine things. It is one job with 9,000 pages in it, and there is no number of seats that makes that cheap, because the cost is in the volume, not the headcount.
The 2026 question for an agency is no longer whether to use AI, but which parts of the year you rent it by the person and which parts you own outright.
Take a book of 180 groups, roughly two-thirds of them on a 1 January effective date, which matches the shape of most small and mid-market books. Each renewal packet is a carrier renewal letter, the current summaries of benefits and coverage at one per plan option across three to six options, the new ones to compare them against, the rate sheet by tier, the census, and on level-funded and self-funded groups a claims and utilisation report with the large-claimant detail on the back pages.
Call it 55 pages a group as a conservative average. That is 9,900 pages in eleven weeks, all of it needing the same three questions asked of it: what changed in the rate by tier, what changed in the benefit design, and what in the claims history explains the ask. Today an analyst does that group by group, and the answer to "which groups got a network change buried in the plan document" is generally found in January, by an angry employee at a doctor's office.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for insurance agency in your browser — 60 seconds, no signup.
The rest of the calendar has its own piles. Form 5500 and the Schedule A data pulled from carriers in July. The prescription drug reporting due on 1 June. The gag clause attestation in December. Creditable coverage notices out to members by mid-October. None of those are hard; all of them are volume.
flowchart TD
A["Renewal packet lands from the carrier"] --> B["Split into rate sheet, SBCs, plan document, claims report"]
B --> C["Read rate change by tier"]
B --> D["Compare new SBC against last year's"]
B --> E["Pull large-claimant trend from the report"]
C --> F["One-page renewal brief for the account executive"]
D --> F
E --> F
F --> G["Account executive checks it against the carrier rep's email"]
In 2024 and 2025 the free-to-download models were visibly the second team. You could run one on your own hardware, and you would notice. That argument closed in 2026. Moonshot AI released Kimi K3, the largest open model in the world, and the open tier as a whole caught up enough that for reading documents and pulling figures out of them the difference stopped being the thing that decides your purchase.
Open, here, means one specific thing to an agency principal: you can download the whole model and run it on machines you own, in your own office or your own rented rack, and no usage meter runs. The bill becomes electricity and hardware instead of a per-person subscription or a per-page charge. It only lights up a fraction of itself for any one job, which is why something that large can run at all outside a data centre the size of a warehouse.
That does not make it the right answer for everyone. It makes it a real second option, where in 2025 it was mostly a hobby.
For most trades this is a pure cost comparison. For benefits it is not, because of what is in the documents. A self-funded group's monthly claims report has diagnosis categories and large-claimant detail. Evidence of insurability forms are health questionnaires. Disability paperwork is medical records with a cover sheet. Every one of those is protected health information, and every vendor that touches it is a business associate who needs a signed agreement and a straight answer about where the data sits.
Most major AI vendors will sign a business associate agreement on their business tiers now. Read the plan you are actually on, not the marketing page. And be ready for the other question: several large employers' security questionnaires in 2026 ask specifically where member data is processed and whether it is used to improve anyone's model. If your answer is a floor plan and a serial number, that questionnaire takes an afternoon. If your answer is a chain of three vendors, it takes six weeks and a lawyer.
That is the honest reason a benefits agency might own a machine when a landscaping company with the same headcount never would.
The hardware is the easy part — a server with a couple of serious graphics cards, sitting in the closet with the phone system, or the rented equivalent. Budget $22,000 once, plus power. The part that gets underestimated is that a machine has an owner, and a 25-person agency does not have that person on the org chart. There is no systems administrator between the compliance manager and the enrollment specialist.
In practice this means your managed IT provider adds it to the monthly contract: patching, backups, monitoring, and being reachable at 6am on the Monday open enrollment opens. Call that $900 a month. Then one person inside the agency — usually the benefits analyst, occasionally a sharp account executive — owns what it is asked to do and checks its output, roughly four hours a week through the fourth quarter and an hour a week the rest of the year.
Still reading? Stop comparing — try CallSphere live.
See the insurance agency AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
If nobody at your agency wants that job, the answer is seats. That is not a failure; it is a correct reading of your own bench.
Assumptions, all illustrative: 25 staff, 180 groups, 9,900 renewal pages in the fourth quarter plus roughly 6,000 pages across the rest of the year, seats at $32 per person per month for the 16 people who would actually use one, and bulk document work priced on the hosted side at the going metered rate for heavy reading.
| Three-year cost | Seats only | Seats plus own machine |
|---|---|---|
| Seats | $18,432 (16 seats) | $8,064 (7 seats) |
| Metered bulk document reading | $14,400 | $0 |
| Hardware | $0 | $22,000 |
| Managed IT add-on | $0 | $32,400 |
| Internal owner's time (150 hours/yr at $41) | $0 | $18,450 |
| Total | $32,832 | $80,914 |
Read that honestly: at 25 people and 180 groups, seats win on cost, and it is not close. Owning the machine starts to make sense at roughly 60 staff and 600 groups, when the metered document reading is the biggest line rather than the smallest, or earlier if a large self-funded client's security requirements make it a condition of keeping the account. Anyone selling you a server at 25 people is selling you their weekend project.
Reading a plan document is not the same as interpreting one. When the summary of benefits and coverage says one thing and the certificate of coverage says another — which happens most often on infertility, gender-affirming care and out-of-network emergency language — the answer is a phone call to the carrier's underwriter and a written confirmation, not a comparison table. Nothing you run, rented or owned, should be the last word there.
The renewal recommendation itself stays human for a different reason. The client's answer is never only about the rate; it is about whether they can survive another year of employee complaints after last year's network change, and whether the CFO will forgive a second disruption. That lives in a relationship, and the account executive who has it should be spending the hours the machine gave back on that conversation, not on page 41.
No. Buy seats for the people who write and quote, get the business associate agreement signed, and revisit in two years. The open-model option is a real one at scale and a distraction below it.
Yes, and most agencies that go this way end up doing exactly that: the hosted tools for daily writing and client-facing work, the owned machine for the bulk reading of anything containing claims data. Splitting on that line — does this document contain member health information — is easier to teach a service team than any rule based on which tool is better.
You download it. That is genuinely the answer, and it is the strongest argument for owning: the hardware is not tied to one model. What you cannot do is escape the next hardware generation, so treat the box as a three-year asset and budget accordingly.
Carriers care about the data agreements attached to whatever holds their claims files, and increasingly they ask. Have the answer written down before the question arrives during a stop-loss renewal.
One thing worth separating out: the volume of documents is a fourth-quarter problem, but the volume of phone calls is a first-quarter one, and they need different answers. CallSphere builds AI voice and chat agents that pick up the service line and the web chat, take the member and group details, book the call-back and pass the note to the right account manager. It does not read your plan documents. It handles the calls those documents generate the week after they land.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Open models closed the gap in 2026. A worked buy-versus-rent comparison for a 58-employee credit union, including the staffing cost nobody budgets for.
400 estoppels and 1,900 inspection photos a month change the math on AI licensing for community association managers. A worked breakeven, plus the limits.
When per-seat AI pricing stops making sense for a host agency or tour operator, what running an open model actually costs in people, and the data case for it.
Kimi K3 and the 2026 open tier let a building materials dealer host instead of license per seat. Worked cost comparison for 4,000 January cost changes.
A 1/1 commercial submission packet costs an account manager nine hours, eight of them gathering. In 2026 you hand over the goal and review the finished packet.
A propane marketer has few desks and enormous record volume. The seat math, the per-use math, and what a box in the closet really costs in people, not dollars.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI