By Sagar Shankaran, Founder of CallSphere
How a community paper builds a 120-call test set from its own obituary intakes, runs an AI agent in shadow mode, and widens its authority one rung at a time.
Key takeaways
What happens the first time your AI agent spells a dead woman's name wrong?
Not hypothetically. Concretely: it takes the obituary by phone from a funeral director at 4:40 on a Friday, hears "Kathryn," sets "Katherine," and 3,900 copies go out Saturday morning with the wrong name on a woman whose family has taken this paper since 1961. You run a correction. You refund the obit. You lose the funeral home, which sends you 60 to 90 paid obituaries a year at $180 to $400 apiece, and you do not get it back, because the funeral director has to face that family at the visitation and you do not.
That single scenario is why most publishers who tried an answering agent in 2024 turned it off inside a month. It is also the exact problem the 2026 tooling was built for. Agent evaluation is the practice of scoring an AI agent against a set of your own real past cases, with a known correct answer for each one, before you let it speak to a customer — and then watching, case by case, what it actually did rather than whether it sounded confident.
Every paper has this material already. Your phone system has recordings or at minimum a call log; your obituary folder has the finished, published, family-approved text of every obit you ran; your billing system has what each one cost. Those three things, joined together, are a test set.
Pull 120 calls from the past year, weighted the way your week actually looks. For a weekly with a Tuesday noon obit deadline, that is roughly: 45 obituary intakes from funeral homes, 25 missed-delivery complaints, 15 vacation stops and restarts, 12 classified line ads, 10 calendar and community-submission calls, 8 legal notice deadline questions, and 5 calls that were something else entirely — a reader who wanted 1987 microfilm, somebody looking for the sports editor, a wrong number that turned into a news tip.
Then build the answer key. For an obituary intake, it is eleven fields: legal name and exact spelling, age, date of death, town of residence, funeral home and its billing account, visitation date, time and location, service date and time, preceded-in-death and survived-by names, photo yes or no, and the word count with the price it produces. That is not a technical exercise. That is your obituary clerk with a legal pad and two afternoons.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent in your browser — 60 seconds, no signup.
Shadow mode means the agent takes the call in parallel and produces its answer, but a human still sends the result. Nobody outside the building knows it is running. For a week or two, every call your front counter takes also gets handled invisibly, and at the end of each day you have two versions to compare: what your clerk wrote down and what the agent wrote down.
This is where the 2026 tooling earns its keep. It is not just a pass-or-fail score. You can see the whole sequence — which field it asked about first, where it accepted an answer it should have questioned, the exact moment it took "Kathryn with a K-A-T-H" and wrote "Katherine." That step-by-step view is what turns a vague sense that "it's pretty good" into a specific fix. The Claude Enterprise governance update on 2 July 2026 added the other half of what a small business needs here: a cost and usage dashboard with spend limits and alerts at 75 and 90 percent, so a runaway agent shows up as an alert rather than as a surprise bill.
flowchart TD
A["Pull 120 recorded intake calls"] --> B["Build answer key: 11 fields per obituary"]
B --> C["Agent runs in shadow, clerk still sends"]
C --> D{"All 11 fields correct on 118 of 120?"}
D -->|No| E["Fix the misses, add failed calls to the set"]
E --> C
D -->|Yes| F["Live for after-hours obituary intake only"]
F --> G["Widen to daytime calls, then to price quotes"]
G --> D
The mistake is treating the go-live as a switch. It is a ladder, and each rung has its own bar to clear.
Rung one: after-hours message taking only. The agent answers at 6:15 p.m., collects the caller's name, number and reason, and emails the front counter. Nothing it says commits you to anything. Rung two: missed delivery and vacation stops — low stakes, high volume, and your district manager gets a clean list at 7 a.m. instead of nine voicemails. Rung three: obituary intake with a mandatory proof, meaning the agent takes the details and immediately emails the funeral home a proof to approve before anything is set. Rung four: quoting the price and confirming the deadline. Rung five, which many papers should never climb, is committing to a run date without a human touching it.
Each rung needs the test set rerun, because widening authority changes what a mistake costs. An agent that is 96 percent accurate is fine at rung one and a liability at rung four.
Assume 26 paid obituaries a week at an average of $265, which is about $344,000 a year and, at most weeklies, one of the top three revenue lines. Assume your current human error rate on intake — a wrong middle initial, a visitation time off by an hour, a missing survivor — is 2 percent, because you have been counting and it is.
| Line | Assumption | Result |
|---|---|---|
| Obituaries per year | 26 per week × 50 | 1,300 |
| Errors at 2% | current human intake | 26 per year |
| Direct cost per error | refund + correction space + 2 hrs staff | $395 |
| Annual direct cost today | 26 × $395 | $10,270 |
| Agent error rate after evaluation | tested at 1.2% with mandatory proof | 15.6 per year |
| Annual direct cost after | 15.6 × $395 | $6,162 |
| Difference | $4,108 |
Notice what this table is not claiming. It does not claim the agent is flawless; it claims a tested agent with a forced proof step beats an interrupted human at 4:40 on a Friday. And notice the line that is missing, because it cannot be honestly estimated: the funeral home relationship. One lost account at 70 obituaries a year is $18,550, which dwarfs everything in the table. That is the number the test set is really protecting.
Still reading? Stop comparing — try CallSphere live.
CallSphere ships complete AI voice agents per industry — 14 tools for healthcare, 10 agents for real estate, 4 specialists for salons. See how it actually handles a call before you book a demo.
Family-submitted obituaries at the front counter stay human. A widow standing at your window with a handwritten page is not a workflow. The death of a public figure — a sitting mayor, a longtime coach, a former publisher — goes straight to the editor, because that is a news story with an obituary attached, not an obituary. Anything where the caller is crying goes to a person, and the agent should be built to hand off the second it hears it.
On the business side, keep a human on deadline commitments. Only your paginator and your press foreman know whether an obit that lands at 11:50 on Tuesday can still make the book. And keep a human on credit — an agent that agrees to run $400 of obituary for a funeral home 90 days past due has cost you more than it saved.
Full test sets are how this stalls. Pull twenty obituary calls, write the answer key for those twenty, and run any of the current agents against them cold, with no tuning at all. You will get a number within an afternoon. If it gets 14 of 20 completely right, you know the shape of the work ahead. If it gets 19, you have a much shorter road than you thought. Either way you now have a measurement instead of an opinion, which is the entire difference between 2026 and the version of this you tried two years ago.
For twenty calls with an eleven-field answer key, an afternoon. For the full 120, budget two to three days of your obituary clerk's time spread over a couple of weeks. It is the least glamorous work in this entire subject and it is the only part that determines whether the rollout succeeds.
Check your state's recording law and whatever your phone greeting already says — if you are recording for quality purposes you likely have what you need, but a dozen states require all-party consent and your greeting has to reflect it. Separately, California SB 53 and Texas TRAIGA both took effect 1 January 2026, and several other states have their own AI statutes; federal preemption is unsettled as of July 2026, so your state's rules still bind. Ask your attorney once, in writing, and keep the answer.
It depends entirely on the rung. For after-hours message taking, if it captures name, callback number and reason on 19 of 20, go. For anything that quotes a price or commits to a run date, you want essentially no misses on your test set and a mandatory proof step behind it, because the errors that survive to print are the expensive ones.
That gap is real and it usually means your test set is too clean — you pulled the calls where the funeral director was organized. Deliberately include your worst recordings: the cell phone in a parking lot, the caller with a heavy accent, the one where three people talk at once. Those are the calls that decide it.
Everything above assumes you have something answering the line in the first place, and that it can be measured. CallSphere builds AI voice and chat agents that answer a business phone line and website chat, take details, book appointments and capture leads around the clock — which for a paper means the 6:15 p.m. obituary call and the 7 a.m. missed-delivery call both get picked up. Run it the way this piece describes: shadow first, your own past calls as the test, and authority widened one rung at a time.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
The 200-call test set a treatment program should build from its own recordings, the five things to score, and how to widen an agent's authority safely.
Newsprint is ordered nine weeks ahead and the draw is a guess. How 2026 forecasting models cut returns, catch sellouts and set renewal prices by cohort.
Why an AI model that “improves” a foreclosure notice costs a community weekly a republication and an affidavit — and what a newsroom-tuned model does instead.
Fourteen box scores, a faxed stat sheet, a penciled scorebook photo. How 2026 document readers cut keying at the sports desk and where they still get names wrong.
How landscape operators split the spring call flood: a fast cheap model for skips and balances, the strong one for brown-out, chemical and cancellation calls.
Before an AI answers your PT clinic's intake line, build a 150-call answer key from last year's schedule. What to sample, how to score, where humans stay.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI