By Sagar Shankaran, Founder of CallSphere
Set per-physician AI caps, a practice ceiling and 75%/90% alerts after the 2 July 2026 governance update - with the cost-per-note math a GI group can check.
Key takeaways
In February, the AI charge on the group's card statement was $611: a pilot, two of the nine gastroenterologists trying ambient note drafting in the exam room on Barrett's surveillance follow-ups and new hepatology consults. By June the same line read $4,812. Nothing had been signed. It had spread the way useful things spread in a busy practice: one physician telling another in the hallway between the 10:15 and the 10:30.
The practice administrator did what any administrator does with a number that grew eight-fold in four months: she tried to break it down by person, so she could tell the managing partner which doctors were driving it. She could not. Until the start of July, the tools told you the total and nothing else — one lump, with no name attached to any of it.
An AI spend cap is a hard monthly dollar ceiling set for each person and for the practice as a whole, with warnings sent at 75% and 90% of that ceiling, so the bill becomes something you approve in advance instead of something you read about on the first of the month. That is the piece that arrived on 2 July 2026, and it is the reason this stopped being a governance conversation and started being a budget one.
The $4,812 was not one thing. It was at least four jobs of wildly different value:
Four different jobs, four different values, one undifferentiated number. Without a breakdown, an owner staring at a frightening total does the one thing guaranteed to be wrong: cuts across the board.
On 2 July 2026, the Claude Enterprise governance update added usage and cost analytics, spend limits at both the organisation level and the individual user level, automatic alerts when a user or the organisation crosses 75% and then 90% of that limit, and the ability to set which model people get by default and what each person is entitled to turn on. There are also reporting tools so your bookkeeper can pull the numbers into the same spreadsheet as everything else.
Compare that with 2025, when you could see a total after the fact and cancel seats. That was the whole dial. Now the administrator sets a ceiling per person before the month starts, is told on the 18th that the hepatologist is at 75%, and can decide — while there is still time — whether to raise his ceiling or move his portal replies to a cheaper default. The bill cannot exceed the number she typed.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for healthcare in your browser — 60 seconds, no signup.
Operationally that moves AI out of the category of "vendor charge we argue about later" and into the category a medical practice already knows how to manage: a budgeted line with an owner, a ceiling, and a monthly review.
flowchart TD
A["Practice sets July ceiling at $3,400"] --> B["Per-person caps loaded: 9 MDs, 2 billers, 2 front office"]
B --> C{"Anyone past 75% before the 20th?"}
C -->|No| D["Month closes; administrator reviews cost per physician"]
C -->|Yes| E["Alert lands with the practice administrator, not the doctor"]
E --> F{"Is the spend clinic notes or appeal letters?"}
F -->|Clinic notes| G["Raise that physician's cap; fund it from the transcription line"]
F -->|Appeal letters| H["Move billing to the cheaper default model"]
The method is unglamorous. Take one clean month of actual per-person usage — now that you can see it — and set each person's cap at roughly 40% above what they used. Then set the practice ceiling below the sum of the individual caps, on purpose, because not everyone spikes in the same month.
Two rules matter more than the numbers. First, the alerts go to the practice administrator, not to the physician. A gastroenterologist part-way through an eight-case endoscopy block should never be reading a spending notification. Second, the entitlements follow the work, not the org chart: the billing office drafting a timely-filing appeal does not need the most expensive model, while the physician dictating a complicated inflammatory bowel disease visit with three prior biologics in the history does.
And put the review where the practice already looks: the AI line belongs on the monthly financials directly underneath the transcription line it is replacing, so the managing partner sees both numbers in one glance.
Assumptions, illustrative, and your practice will differ: nine physicians across two clinic sites and an endoscopy center; roughly 3,400 signed office-visit notes a month; June's per-person usage as read off the new dashboard; 40% headroom on individual caps.
| Group | People | June actual, each | Cap set, each | Sum of caps |
|---|---|---|---|---|
| Physicians | 9 | $312 | $440 | $3,960 |
| Billing and appeals | 2 | $208 | $290 | $580 |
| Front office and referrals | 2 | $96 | $150 | $300 |
| Practice ceiling | — | $4,812 total | — | $3,400 |
The individual caps add up to $4,840. The practice ceiling is set at $3,400, which is 30% lower. That is not an accounting error — it is the point. In June only three of the nine physicians were above $400; the average was dragged upward by two heavy users. Individual caps stay generous enough that nobody gets blocked in the middle of a Thursday clinic, while the practice ceiling protects the group if several people spike in the same month.
Now the value check the managing partner wants. At a $3,400 ceiling across 3,400 notes, the practice pays about $1.00 per signed note; the transcription service it replaced ran closer to $1.75 at the same volume, so the ceiling covers itself on that comparison alone. The after-hours charting time each physician gets back — 45 minutes per clinic day, three clinic days a week, roughly 9.7 hours a month — never appears on any invoice, and it is the part that keeps partners from calling the recruiter who emails them every Tuesday.
The most expensive mistake an owner can make with a spend cap is to cut the biggest user first. In a GI group that is usually the highest-volume physician, and saving $180 a month by capping her note drafting so she charts for an extra hour at night is a trade no practice should take. Cut the users whose spend has no matching hour saved — a completely different list, and now a visible one.
Still reading? Stop comparing — try CallSphere live.
See the healthcare AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
A cap tells you what something cost, not whether it was right. Those are separate reviews, and the second one protects a license: each physician should genuinely read, not merely sign, a sample of drafted notes every month, and the billing manager should spot-check that appeal letters quote the payer's own written policy rather than a plausible paraphrase of it.
A cap also cannot catch the real leak: a staff member pasting a patient's chart into a free personal account on their phone because the practice tool felt slow that morning. No dashboard sees that. That is a written policy, a signed business associate agreement, and five honest minutes at the staff meeting.
One seasonal warning: set January differently. Its first two weeks bring the deductible reset, every plan change from open enrollment, and a wave of re-verification work. Alerts calibrated to a quiet October will fire all month on normal January volume, and staff will learn to ignore them.
Both, and they should not add up. Individual caps stop one person running away with the budget; the practice ceiling stops a bad month. Set the individual caps loose enough that nobody is blocked during clinic, and the ceiling tight enough that you would happily see that number on the statement every month for a year.
They go back to typing, which is precisely the disruption you do not want. That is why the 75% and 90% alerts matter more than the ceiling itself — they give the administrator a week of warning. Write a standing rule that a physician's cap can be raised the same day, without a partners' meeting, and reviewed afterwards.
Yes, if it bills under its own tax identification number, as most ambulatory surgery centers do. Separate ceilings keep the cost allocation defensible when partners own different percentages of each entity.
Put the AI line and the cost it was meant to displace side by side for three consecutive months — transcription, scribe hours, billing-office overtime. If the old line has not fallen, you bought an addition rather than a replacement. That can still be a good purchase, but it should be an explicit one.
Open the usage dashboard, sort by person, print one page, and circle the three names you did not expect. Then set every individual cap at 40% above that person's actual usage, set the practice ceiling 30% below the sum, and take it to the next partners' meeting as a budgeted line. Ninety minutes of an administrator's time, and a recurring argument becomes a number.
One line on that dashboard will be the phone and web chat that answers when the office is closed — the prep question at 9:40 p.m., the new referral calling at 6:40 because the primary care office just faxed it over. CallSphere builds the voice and chat agents that answer those calls, book the appointment and capture the lead around the clock. Like every other line on that page, it deserves a name against it, a ceiling, and a monthly review next to the answering-service invoice it replaced.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
AI spend in retail follows floor traffic, not the calendar. Set seasonal caps, per-user limits and 75% alerts without killing your December chat agent.
How a precision machining owner budgets, caps and reviews AI spend after the 2 July 2026 Claude Enterprise governance update, with a cost-per-quote example.
Spend limits, 75% and 90% alerts and entitlements arrived in July 2026. A per-role AI budget for an imaging center, plus cost per authorized study.
A 40-claim DME probe takes 16 working days one at a time and finds the gaps too late. Split four ways, the records requests go out on day two of forty-five.
Consignee changes, bank details, hold releases and DEA calls: how a contract manufacturer verifies the caller when the voice itself proves nothing in 2026.
Vietnamese, Spanish and Mandarin-speaking dental front offices order the minimum and ask nothing. Live translation in 2026 changes the lunch-hour call.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI