By Sagar Shankaran, Founder of CallSphere
AI costs fell roughly 10x into 2026, so testing every invoice line against every lease exclusion is now cheap. Worked arithmetic for a nine-building portfolio.
Key takeaways
You looked at this in 2024 and walked away, and you were right to. Somebody proposed running an AI over every vendor invoice to check it against the recovery language in every affected lease, you did the arithmetic on nineteen thousand invoice lines against a hundred and eighty leases, and the number came back somewhere north of what a part-time CAM analyst costs. So you did what everyone in this business does: you kept sampling, you kept your fingers crossed through reconciliation season, and you set aside a reserve for the audit letters.
The arithmetic has moved. Frontier AI now costs roughly a tenth of what it cost in 2025, capable models are priced at a level where reading a document is a rounding error, and running high-volume routine work on your own hardware is about 90% cheaper again than sending it to the cloud. The check you priced and rejected is now cheaper than the certified mail you use to send the true-up statements.
Here is the work in plain terms. Your property accountant codes an invoice — the January snow removal bill from the landscaping contractor, $11,400 across three buildings. It gets a general ledger account and a recoverable-or-not flag. That coding is made once, by one person, based on how this building has always been coded.
What is not done is the per-tenant check. Because the truth is that $11,400 does not have one answer. In Building C, the tenant in Suite 300 negotiated Exhibit C language excluding snow removal above a stated annual amount. Two suites down, a national tenant capped controllable expenses at 5% cumulative and compounded, so once you cross the cap it does not matter how correct the coding is. In Building F, the ground-floor restaurant pays a separate pro-rata on the parking field only. And there is a tenant whose lease excludes any management fee charged on top of a capital repair, which changes whether the administrative fee on this invoice is billable at all.
The definition to hold onto: a per-transaction check means every single invoice line is tested against the actual recovery language in every lease it touches, at the moment it is coded, instead of once a year at reconciliation on a sample.
Nobody does that by hand. A property accountant at a mid-sized owner processes several hundred invoices a month across nine buildings. Reading 180 leases per invoice is not a workflow, it is a fantasy. So the trade sampled, and reconciliation season in the first quarter became the moment the mistakes surfaced — usually via a tenant's lease administration department, which does this professionally and gets paid a share of what it recovers.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for real estate in your browser — 60 seconds, no signup.
In 2025, having a careful model read an invoice with the relevant clauses from every affected lease attached, and return a recoverability opinion with citations, would have run you a few cents per line at best. Nineteen thousand lines, plus the lease text going in every time, and you were looking at a four-figure annual bill before anyone had built anything — against a benefit you could not size in advance. That is a reasonable thing to decline.
Through 2026 the price of the same work fell by roughly a factor of ten. Strong models are now priced at about $2 for a million words of text handled, and the routine part of this job — matching an invoice description to a lease exclusion list you have already extracted once — can run on a cheap fast model, or on hardware you already own, at around a tenth of even that. The lease language only has to be read carefully once per lease per year, not once per invoice.
flowchart TD
A["Vendor invoice arrives in AP"] --> B["Coded to GL account and building"]
B --> C["Checked against the recoverable pool rules"]
C --> D{"Capital or operating?"}
D -->|Capital| E["Held out; amortisation question to controller"]
D -->|Operating| F["Tested against each affected tenant exclusion and cap"]
F --> G{"Any tenant-specific exception hit?"}
G -->|Yes| H["Flagged with lease clause cited for accountant review"]
G -->|No| I["Posted to the recovery ledger, ready for the true-up"]
The invoices come in through the day. Each one gets read as it is coded, against a set of rules that were built once, at the start of the year, by extracting the recovery language out of every lease in the portfolio. Most invoices pass silently. The elevator maintenance contract, the day porter, the fire alarm monitoring — nobody negotiated a carve-out on those and nothing happens.
Three or four times a week, something stops. The parking lot resealing invoice comes in coded as a repair; the check flags it against four leases that define resurfacing as capital, with the clause and page number attached. The roof consultant's $6,800 report gets flagged because two leases exclude consultants' fees from the CAM pool by name. The security guard company's invoice includes a fuel surcharge, and one lease excludes surcharges of any kind.
The property accountant looks at each flag, spends two minutes, and either agrees or overrides with a note. By the time reconciliation season arrives, the recovery ledger has already been argued with, invoice by invoice, all year. The reconciliation stops being a discovery exercise and becomes an arithmetic one — which is the whole point.
Assume nine buildings, 180 tenants, $4.8 million of annual recoverable operating expense, about 7,400 payable invoices a year averaging 2.6 lines each — call it 19,200 lines. Assume the routine per-line check costs a quarter of a cent, and that about 6% of lines get flagged and pushed to a careful model at four cents to produce a citation. Assume one-time lease extraction of 180 leases at $1.20 each.
| Item | Volume | Unit | Cost |
|---|---|---|---|
| Routine per-line check | 19,200 | $0.0025 | $48.00 |
| Flagged lines, careful review with citation | 1,152 | $0.04 | $46.08 |
| One-time lease rule extraction | 180 | $1.20 | $216.00 |
| Property accountant review, 2 min per flag | 38 hrs | $41/hr | $1,558.00 |
| Total first year | — | — | $1,868.08 |
Now the other side, and be honest that this is an illustration rather than a promise. Suppose the check corrects errors on just 0.7% of the recoverable pool. On $4.8 million that is $33,600 — some of it over-recovery you would have refunded with an audit fee attached, some of it under-recovery you simply never billed and never knew about. Under-recovery is the quiet one. Nobody writes you a letter about the $9,000 you forgot to charge.
Prove it on your own portfolio without committing to anything: take last year's completed reconciliation for one building, run the check retroactively against the invoices you already posted, and count the disagreements. If it finds nothing, you have a clean shop and you have learned that for the price of a slow afternoon.
Still reading? Stop comparing — try CallSphere live.
See the real estate AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
It does not settle a gross-up dispute. If your building ran at 78% occupancy and you grossed variable expenses up to 95%, the argument with the tenant's auditor is about methodology and about what the lease permits, and no per-invoice check touches it. That is a conversation between your controller, the tenant's lease administrator, and sometimes counsel.
It does not decide capital versus expense. It can flag that a $6,800 roof item looks capital under four leases, but whether resurfacing 40% of a parking field is a repair or a capital replacement — and over what useful life you amortise it, at what interest rate the lease permits — is a judgement your controller makes and defends. Let the check raise the question; do not let it answer it.
And it will not save a portfolio where the underlying lease abstracts are wrong. Every rule in this system is only as good as the extraction it came from. Spot-check the extracted exclusion list against the actual Exhibit C on your twenty largest tenants by square footage before you trust a single flag.
No, but it changes the meeting. When a national tenant's audit firm arrives with a findings letter, the difference between a defensible position and a bad week is whether you can show why each contested line was coded the way it was. A year of flags, overrides and clause citations is exactly that record.
Not really. The check sits at the point of coding, wherever that happens, and writes a flag back. The one thing to insist on is that the flag and the override note end up attached to the invoice record itself, not in a side spreadsheet, or the record is worthless twelve months later.
It is an illustration, not a benchmark, and you should treat it that way. The reason to run the retroactive test on one building is precisely to replace that number with your own. Portfolios with heavily negotiated national retail leases usually find more; single-tenant industrial with clean absolute net leases usually finds almost nothing.
Then you have a bigger problem than CAM coding, and you have just found it, which is worth something. Start with the leases you have, and put the missing ones on the abstracting list before reconciliation season.
The first quarter is also when your phone changes character. True-up statements go out, and for three weeks the management office line is tenants asking why their monthly estimate went from $3,900 to $4,340 — the same six questions, over and over, from people who are annoyed. CallSphere builds AI voice and chat agents that answer that line 24/7, take the building and suite, answer the routine statement questions, and route anything that sounds like a dispute straight to your property manager, so reconciliation season stops eating your team's afternoons.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Per-line payroll review was uneconomic in 2025. After a roughly 10x price drop, checking every register line before funding pays for itself on class codes.
Model routing for a management office: which tenant emails a fast model can answer, which need a lease read, and the escalation triggers that protect you.
Why multi-unit franchise operators sampled the exception report, what the 2026 collapse in AI cost changes, and a worked example on 4.6 million tickets a year.
Why per-note golden-thread checking became economic in 2026, what the morning exception list looks like, and the line you must not let the software cross.
A full-file coding check reads every primary care encounter before the claim goes out. What it costs, what it finds, and where a coder still has to decide.
Which property management roles change shape when AI drafts the abstract and the recovery schedule, what to teach in week one, and what leaves the job posting.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI