By Sagar Shankaran, Founder of CallSphere
A full-file coding check reads every primary care encounter before the claim goes out. What it costs, what it finds, and where a coder still has to decide.
Key takeaways
Three thousand two hundred. That is roughly how many closed encounters come out of a three-clinician family medicine practice in a quarter — two physicians and a nurse practitioner, twenty patients a day each, four days of clinic a week. Now count how many of those charts anyone actually looked at a second time. In most practices the answer is ten a month, because that is what an outside coding review costs when a credentialed coder is reading at forty dollars an hour. Ten out of thirty-two hundred. No quality committee would accept that sample on a clinical measure, but it has been the standard for revenue integrity for twenty years.
The reason was never that owners did not care. It was arithmetic. Reading a chart carefully — the note, the assessment and plan, the diagnosis codes, the claim lines, the modifiers — took a human eight to twelve minutes and cost real money. Checking all of them cost more than the errors were worth. That stopped being true this year.
A coding sample answers one question: does this practice have a systemic habit. It does not answer the question owners actually have, which is which claims going out this week are wrong. The two are not the same. A single physician who has drifted into billing 99213 for visits that clearly document moderate-complexity decision-making will not show up in ten random charts unless you get lucky. Neither will the annual wellness visit that was performed, documented, and then never billed because the medical assistant closed the encounter before the G0439 line was added.
Here is the definition worth keeping: a full-file coding check is a review of every closed encounter — note, diagnosis codes, and claim lines — before the claim leaves the practice, instead of a monthly sample of a handful of charts. It is not new as an idea. Hospital systems have had rules engines bolted onto their billing for a decade. What is new is that a small independent practice can now afford one that actually reads the narrative note rather than pattern-matching the code fields.
Today the workflow is this. The clinician closes the encounter in eClinicalWorks, athenaOne, Elation or NextGen. The claim drops to the billing work queue, and your billing specialist runs it through the clearinghouse scrubber at Availity or Waystar. That scrubber catches format problems: a missing referring provider, an invalid place of service, a truncated diagnosis code. It does not read the note. It has no opinion about the code that should have been on the claim and is not.
So the claim goes out. Six weeks later the remittance comes back with CO-11, diagnosis inconsistent with the procedure, or CO-197, precertification absent, or nothing at all — because the most expensive error in primary care is not a denial. It is the claim that pays exactly as submitted, at a level below the work that was documented, forever, without complaint. Nobody appeals money they never asked for.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for healthcare in your browser — 60 seconds, no signup.
The workaround everyone pretends is fine: the practice manager pulls a monthly report showing each clinician's mix of 99212 through 99215, sees that Dr. B runs 68% level-three visits against a group average of 44%, and mentions it at the Thursday huddle. That is management by aggregate. It has never once corrected a specific claim.
Between 2025 and 2026 the price of running a strong model fell roughly tenfold. In plain terms: having a capable system read a million words of your own clinical notes now costs a couple of dollars, and running the same work on hardware sitting in your own office is roughly ninety percent cheaper again for high-volume jobs. A primary care progress note, its problem list, and the claim lines attached to it come to somewhere between 1,200 and 3,000 words. At 2026 prices, having that read carefully and compared against the codes submitted costs a few cents per encounter — and it costs the same on the three-thousandth chart as on the first.
That single change flips the economics. In 2025 you sampled because reviewing everything cost more than it returned. In 2026 the review is the rounding error and the sample is the expensive habit.
flowchart TD
A["Clinician closes encounter in the EHR"] --> B["Claim lands in billing work queue"]
B --> C["Overnight check reads note + codes + modifiers"]
C --> D{"Does the note support the level billed?"}
D -->|Supported| E["Claim releases to clearinghouse"]
D -->|Under-documented| F["Held: query back to clinician"]
D -->|Under-coded or missing line| G["Flag to billing specialist with the note excerpt"]
F --> B
G --> H["Coder decides, corrects, releases"]
The check runs on the day's closed encounters. By 7:15 the billing specialist has a worklist of nine items out of 84 encounters, and each item is one line of plain English with the exact sentence from the note that triggered it.
The specialist works the list in eighteen minutes. Nothing was automatically changed. The clinician gets two queries, not nine.
Stated assumptions, all of which you should replace with your own fee schedule and your own denial report. This is an illustration, not a promise.
| Assumption | Value |
|---|---|
| Closed encounters per quarter | 3,200 |
| Cost to read and check each encounter | $0.03 |
| Total quarterly checking cost | $96 |
| Encounters flagged | 6% = 192 |
| Flags a credentialed coder agrees with | 50% = 96 |
| Of those: level corrections (avg. difference) | 60 × $38 = $2,280 |
| Missed wellness or care-management lines | 22 × $95 = $2,090 |
| Denials prevented (rework at $25, 30% never recovered) | 14 × $61 = $854 |
| Quarterly gross effect | $5,224 |
| Less checking cost and 6 hours of coder time at $40 | −$336 |
| Net per quarter | about $4,888 |
Two things to notice. The checking cost is not the number that matters — the coder's time is, which is why flag precision is the whole ball game. And this is not a growth story. It is money the practice already earned.
Do not let anything change a claim by itself. A system that quietly upgrades levels of service is building you a pattern that a payer audit or a Unified Program Integrity Contractor will read as intent. Every flag goes to a human who signs off. If your billing is outsourced, that human is at the vendor, and you should ask them in writing who reviews and how the decision is recorded.
Still reading? Stop comparing — try CallSphere live.
See the healthcare AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
It is also weak in three specific places. Time-based billing where the time statement is vague ("spent considerable time counseling") — it cannot manufacture what the clinician did not write. Collaborative care codes, where the requirements sit in the care manager's registry activity and not in the visit note at all. And any incident-to arrangement, where the answer depends on who was physically in the suite, which no note reliably records.
One more honest limit: a flag that says "the note under-documents the work performed" is a documentation problem, not a coding win. Chasing those with retroactive addenda is how practices get in trouble. Fix the template going forward instead.
Reviewing your own claims before you submit them is the opposite of audit risk — it is what payers ask practices to do. The risk comes from what you do with the findings. Correcting a claim before submission is routine. Systematically raising levels of service without changing what is documented is not, and volume makes that pattern visible faster, not slower.
Their job is getting submitted claims paid. Almost no percentage-of-collections vendor reads narrative notes to find work you never billed for, because their margin sits in clean submission and follow-up. Ask yours whether they review notes against levels, and on what sample. The answer is usually under twenty a month.
It needs two things: the note text and the claim lines for the same encounter. Every certified system can produce both, whether through its reporting module, a nightly extract, or the standard data export required for certification. It is a boring integration job, and it is where most of the setup effort actually goes.
Run the check backwards over 300 encounters you already closed last quarter. Correct nothing. Just see whether the flags hold up when your coder reads them. Under a third holding up means the setup is wrong and you lost a morning. Two-thirds means you have a number for your board.
Pull your last four remittance advice files and sort denials by reason code. Then pull your level-of-service distribution by clinician. If the second report is flat and the first is dominated by CO-11 and CO-4, your money is leaking through documentation-to-code mismatch, which is exactly what a full-file check catches. If your denials are all CO-197, your problem is prior authorization and this is not your first project.
A note on the other end of the same phone line: most of the corrections above start with a patient interaction that a front desk never captured cleanly — the wellness visit that was scheduled as a sick visit, the discharge callback that missed its window. CallSphere builds AI voice and chat agents that answer practice phone lines and web chat, book appointments, and capture what the caller actually wanted, around the clock. It does not code your claims. It does make the encounter start with the right reason for visit attached, which is where a surprising share of coding cleanup begins.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
How rural primary care clinics in Crescent City, California can answer every patient call 24/7 with AI — after-hours triage, booking, and referral follow-up.
Nampa, Idaho family practices face Treasure Valley growth that outpaces front desks. How an AI answering service absorbs new-patient call volume 24/7.
Estes Park, Colorado family clinics face summer call surges and quiet winters. An AI receptionist scales with Rocky Mountain National Park crowds year-round.
Harrisburg, Pennsylvania family practices miss commuter calls at 8 a.m. and 5 p.m. An AI receptionist answers around the clock and books every state worker.
US small-business AI adoption hit 66% but 70% of owners say staff need training. What changes on an optometry org chart and what week one must now cover.
Per-line payroll review was uneconomic in 2025. After a roughly 10x price drop, checking every register line before funding pays for itself on class codes.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI