By Sagar Shankaran, Founder of CallSphere
Deloitte found 84% of AI investors report positive returns. For P&C carriers the provable process is FNOL intake - here are the four numbers to baseline first.
Key takeaways
"We ran a pilot in 2024, it cost us six months, and the only thing anyone could tell me at the end was that people liked it." That sentence, or a version of it, comes up in almost every claims leadership meeting where AI is on the agenda. The complaint is fair. The pilot usually failed for a reason that had nothing to do with the software.
Deloitte's State of AI in the Enterprise 2026 found 84% of organisations investing in AI report positive returns. The interesting part is not the percentage — it is how consistent the pattern behind it is. Map one messy process. Put a human in the review seat. Prove time saved or errors reduced. Then widen. The 2024 pilots that produced nothing almost all skipped the second word in that sequence: one.
Here is what typically happened. A carrier picked three or four uses at once — claim summarisation, a submission reader, a chat assistant for the service desk — ran them a quarter, then tried to work out whether anything improved. Nobody had recorded the average speed of answer in the week before, or the days from first notice of loss to first adjuster contact, or what share of intakes reached ClaimCenter missing the loss location. Without those, the review meeting is a debate about impressions, and impressions lose to budget every time.
The clean version, worth quoting: an AI return-on-investment case in claims is a before-and-after measurement of one named process, with the baseline captured for at least thirty days before anything is switched on — otherwise you are arguing about feelings in a finance meeting.
Thirty days is not a formality. Claims volume is seasonal and lumpy. If your baseline month contains a hail outbreak and your test month does not, every number moves and none of them mean anything. Baseline a period that includes at least one bad week, or take two baselines and say so.
For a regional property and casualty carrier the answer is almost always first notice of loss intake, and specifically the intake that happens outside business hours or during a catastrophe surge. It is the right choice for four reasons an owner can check without a consultant.
It has volume, so a month produces enough events to mean something. It has a natural clock — the loss happened at a time, the call came in at a time, the adjuster made contact at a time — so cycle time needs no invented metric. It is already reported: your Market Conduct Annual Statement buckets claims by days to close, and your state complaint index reflects how long people waited. And the failure is visible to the policyholder, so improvement shows up in retention, not only in expense.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for insurance agency in your browser — 60 seconds, no signup.
The processes that demo well and measure badly are claim summarisation and coverage-question answering. Both are useful. Neither gives you a number a chief financial officer will accept in October, because "the adjuster read it faster" has no clock attached to it.
flowchart TD
A["Pick one process: after-hours and surge FNOL"] --> B["Log 30 days of baseline before anything is switched on"]
B --> C["Agent answers overflow and the 2 a.m. calls"]
C --> D["Adjuster reviews every intake for the first two weeks"]
D --> E{"Loss-report to first-contact time improved?"}
E -->|No| F["Fix the intake questions, keep the same baseline"]
F --> C
E -->|Yes| G["Widen to one more state and re-baseline"]
G --> B
Only four, and every one of them can be pulled from systems you already run. First, hours from loss report to first adjuster contact, measured as a median rather than an average, because one file that sat for three weeks will drag an average and hide the real picture. Second, the share of calls abandoned before a person picked up, split into business hours and after hours, taken from the phone system rather than from anyone's memory.
Third, the intake completeness rate: the percentage of new claims that arrive in ClaimCenter with date of loss, loss location, cause of loss and a working callback number all present. Your claims managers will guess this at 90% and it will come back at 70%, and the missing pieces are what generate the callback that costs another eleven minutes. Fourth, the reopen rate at 60 days, which is your check against the obvious risk that everything got faster and worse at the same time.
Write those four on one page with the dates the baseline covers, get the claims vice president to sign it, and put it in a drawer. That page is what makes the October conversation short.
Suppose the baseline month contains an ordinary three weeks and then a Thursday evening supercell across the Dallas–Fort Worth metro. Normal week: 1,050 first notices, median 19 hours to first adjuster contact, 4% abandoned. Hail week: 3,400 first notices in 72 hours against an intake team sized for 1,100, median time to first contact stretching to 61 hours, abandonment at 14% on the Friday and 19% on the Saturday morning, and a queue of voicemails nobody clears until Monday.
That contrast is the most valuable thing in the baseline, because the surge is where the case actually lives. Nobody buys an agent to handle a Tuesday in February. They buy it because the alternative on the Friday after a hail event is calling independent adjusters and paying surge rates for people who spend their first day taking intake calls rather than inspecting roofs.
Run the test in the following storm season on one state and one peril. Let the agent take the overflow and the overnight calls — date and time of loss, loss location, cause, damage description, habitability, whether emergency mitigation is already on site, the callback number — and have an adjuster read every intake for the first two weeks. Not a sample. Every one, because that is how you find out the agent never asks whether the water is still running.
Two numbers close the case: cost per first notice handled, and median hours from loss report to first adjuster contact. Assume 4,800 first notices a month across the book, 22% of them arriving outside business hours, and an intake cost of $9.40 per call today counting the representative, the supervision and the overflow answering service.
| Measure | Baseline (30 days) | After 60 days |
|---|---|---|
| First notices handled per month | 4,800 | 4,800 |
| Cost per intake | $9.40 | $5.10 |
| Monthly intake cost | $45,120 | $24,480 |
| Abandonment, after hours | 14% | 2% |
| Median hours to first adjuster contact | 19 | 6 |
| Intake completeness | 71% | 93% |
| Reopen rate at 60 days | 8.1% | 7.9% |
A monthly saving of $20,640 against setup and running costs of, say, $38,000 in year one gives payback inside the second month of live running. But the line that carries the meeting is not the cost line — it is 19 hours down to 6, because faster first contact is the one lever that reduces both cycle time and the chance a policyholder hires a public adjuster or an attorney before you have spoken to them. Put the reopen rate on the same slide unprompted; showing the metric that could have embarrassed you is what makes the rest believable.
Still reading? Stop comparing — try CallSphere live.
See the insurance agency AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
Coverage decisions do not belong in a time-saved column. If a project starts reporting "coverage questions resolved per hour," someone will optimise it, and the cost of a wrong denial is a bad-faith exposure that dwarfs any efficiency gain. Keep coverage with the adjuster and out of the scorecard.
Reserving is the same. A faster reserve is not a better reserve, and your actuaries will rightly object to any measurement that treats speed as quality. If AI touches reserving at all, measure accuracy against ultimate development, not turnaround.
And be careful measuring anything that varies by protected class or geography without checking how it varies. Colorado's insurance rules already require carriers to test consumer-facing models for unfair discrimination, and several other states are moving the same way, with the federal position on state AI rules still unsettled as of this July. A metric that improves overall while getting worse in one zip code is a market conduct finding waiting to be written.
Thirty days of baseline and sixty days of live running, minimum, and if your book is catastrophe-exposed then both windows need to contain a comparable event. Anything reported at the four-week mark is noise dressed as a result.
Not the person who chose the software. Give the baseline page and the after-numbers to claims operations or to finance, and let the sponsor argue against them. The 2024 pilots that died were mostly measured by their own advocates, which is why nobody upstairs believed the result.
Then you have learned something for the price of one process and one quarter, which is the entire point of choosing one. In practice flat results in intake usually trace to intake questions in the wrong order rather than to the agent itself — fix the script and re-run against the same baseline rather than starting over.
Disclose it. Several states now require it in consumer interactions and more are moving that way, and beyond the rule it is simply better practice on a claims line. Say it in the first sentence and give an obvious route to a person. Carriers that hide it get complaints about the hiding, not about the agent.
Pull the four numbers for the last thirty days. Do not buy anything, do not book a demo, do not tell anyone outside claims operations. If your phone system cannot split after-hours abandonment from business-hours abandonment, that is finding number one, and it is worth a week of somebody's time on its own.
When you do get to the running part, the after-hours phone line is where most carriers start, and it is where CallSphere fits: AI voice and chat agents that answer the claims line at 2 a.m. and during the Friday surge, take a complete first notice, capture the callback details and book the adjuster contact, with every call logged so the four numbers you wrote down are actually measurable afterwards. Start with the measurement, though. The software is the easy part.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Map one messy process and prove it: spray ticket records, the five baseline numbers to capture before you start, and the error rate that ends the debate.
The past-due report is the gym process to baseline before buying AI: five numbers to capture, a worked example on a 1,400-member club, and the honest limits.
How P&C carriers use computer-use agents on MSPRP, state comp and salvage title portals to free recovery clerks and keep the Section 111 penalty clock clean.
How a security guard company proves AI paid for itself: map the open-shift callout, take a 30-day baseline, and track overtime as a share of billed hours.
Pull twelve months of factor deductions, convert to chargeback dollars per $100,000 shipped, and you have a baseline that settles the AI argument in 90 days.
How a benefits agency proves AI paid for itself: reconcile every carrier commission statement, capture five baseline numbers, and count recovered dollars.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI