By Sagar Shankaran, Founder of CallSphere
Build a test set from your own closed files, score misses apart from false alarms, and widen an agent's authority in three gates before it drafts Schedule B.
Key takeaways
Run a name search on a seller called John A. Miller in a county of any size and you get back somewhere between forty and two hundred hits. Judgments, state tax liens, child support, an old federal tax lien, three UCC filings, a divorce decree, two probate matters. Your title examiner clears almost all of them in an afternoon, because almost none of them are your John A. Miller. The share of name hits that attach to your actual seller is small — call it under ten in a hundred — and the entire value of an experienced examiner is telling which ten without ordering a certified copy of every one.
That is the job people now want to hand to an AI agent. It is also the job where being wrong is not an inconvenience, it is a claim. So before you let one anywhere near a live commitment, build the thing that decides whether it is good enough: a test set out of your own closed files.
Watch a good examiner for twenty minutes and you will see a sequence that almost nobody has written down. She reads the hit. She compares the middle initial and the suffix. She checks the address on the instrument against the chain on the tract. She looks at the date — was the judgment entered before or after your seller took title, because that decides whether it attached at all. She checks whether the docket shows a satisfaction. She checks the amount against how plausible it is for this property. If it is a federal tax lien she checks the refiling date and the redemption period. If it is a mechanic's lien she checks whether the statutory period has run. Then she writes two sentences in the file explaining her disposition, and those two sentences are the most valuable text your agency produces.
An agent evaluation in title work means running the agent over files whose outcome you already know, and comparing its disposition to the disposition your licensed examiner actually recorded — not asking it questions and admiring the answers. The difference between those two activities is the difference between a pilot and a demo.
You do not need to build data. You need to pull it. Take the last eighteen months of closed residential files and select roughly 400, weighted the way your actual order flow is weighted — if a fifth of your volume is refinance, a fifth of the test set is refinance. Then deliberately over-sample the ugly ones, because the average file teaches you nothing. Include:
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for real estate in your browser — 60 seconds, no signup.
For each file, the answer key is what your examiner concluded and what ultimately appeared on Schedule B-I and B-II of the issued commitment, plus anything that had to be fixed afterward. That last column is the one that matters most.
flowchart TD
A["Pull 400 closed files with examiner dispositions"] --> B["Agent runs each file in shadow, no output to customers"]
B --> C["Compare agent disposition to examiner disposition"]
C --> D{"Any missed true lien?"}
D -->|Yes| E["Review the step record: which index, which instrument"]
E --> B
D -->|No| F["Widen: agent drafts B-II, examiner signs every file"]
F --> G["Spot-check 10% monthly against the answer key"]
The thing that matured in 2026 is not the models — it is the ability to see, step by step, what the agent actually did to reach a conclusion. For a title agency that means a record you can open next to a file number and read: it pulled the grantor-grantee index for these years, it opened book 4127 page 812, it compared the middle initial, it treated the satisfaction recorded eleven months later as releasing the 2019 mortgage, and it decided the 2021 judgment against a John Miller with no middle initial and a different last-known address was not your seller.
That record is what turns a disagreement into a training conversation. When the agent and the examiner differ, you are not arguing about whether the machine is smart. You are looking at exactly which step it took wrong — and in practice you will find the errors cluster in three places: name variance (married names, Jr. and Sr., an anglicised spelling), timing (a lien that attached before or after the vesting deed), and satisfactions recorded in a different instrument type than usual for that county. Once you know the clusters, you know which files to route away from it entirely.
Two numbers matter and they are not the same number. A false alarm — the agent flags something that is not your seller — costs an examiner ten minutes. A miss — a real lien it clears as not-our-party — costs you a policy claim. Score them separately and set your threshold on the second.
| Assumption (illustrative) | Figure |
| Files in the test set | 400 |
| Files containing at least one true attaching lien | 62 |
| Agent misses (true lien cleared as no-match) | 2 of 62 = 3.2% |
| False alarms across the 400 files | 115 |
| Examiner time per false alarm | 10 minutes |
| Annual residential volume | 1,080 files |
| Expected true-lien files per year | 167 |
| Expected misses per year at 3.2% | about 5 |
| Average cost to cure a missed lien after closing (payoff, legal, staff) | $9,500 |
| Expected annual cost of a 3.2% miss rate, unsupervised | about $47,500 |
| Same agent with examiner review of every flagged and every cleared match | misses caught before issue |
Read that table the right way round. It does not say the agent is bad; a 3.2% miss rate on lien matching is not embarrassing. It says that at your volume, 3.2% unsupervised costs about $47,500 a year, which is more than the examiner hour you were trying to save. So the agent runs as a first pass, the examiner reviews, and the saving comes from the 115 false alarms she now spends two minutes on instead of ten — because the step record shows her immediately why it flagged.
Widen its authority in stages, and write the promotion criteria down before you start so nobody negotiates them later under closing-week pressure.
Still reading? Stop comparing — try CallSphere live.
See the real estate AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
Your 400 files come from your counties, your search vendor and your last eighteen months of market. They will not tell you how the agent behaves in a county you just expanded into with a different indexing convention, or on a product you do not currently write much of, or during a refinance wave when volume triples and the pressure to skip the review is highest. Re-run the test set whenever any of those three change.
They also cannot tell you about the file where the right answer was not in any record — where an examiner picked up the phone, called the clerk in a rural county, and learned that the release had been recorded under the wrong instrument type in 2014 and the index was never corrected. That call has no digital trace and never enters your test set. It is also, on maybe one file a month, the entire reason a closing happens on time. Keep the person who knows to make it.
Build it yourself, and treat a vendor-supplied benchmark as marketing. The whole point is that it comes from your counties, your recording conventions, your underwriter's bulletins and your examiners' judgment. A test set from someone else's files measures somebody else's business.
Then your first project is not AI. Change the note template so every disposition records the reason in one line, and in ninety days you will have both a usable answer key and a materially better file for defending a claim. Agencies that have done this generally find the note discipline pays for itself before the agent ever arrives.
Assembling 400 files and their dispositions is a week of somebody's part-time attention if your closing software exports cleanly. Shadow running is four to six weeks. Call it a quarter before you are drafting, and do not compress it because a busy summer makes the saving look tempting.
Your most experienced examiner, with your escrow manager as the second reader — not the person who is best with computers. The scarce skill here is knowing what the right answer was, not knowing how to run the tool.
One connected point. The same evaluate-before-you-trust logic applies to anything that speaks to your customers, including the phone. CallSphere builds AI voice and chat agents that answer title and escrow lines and web chat, take status questions, and book signings; the sensible way to bring one in is the same three gates — listen to real recordings against real calls first, widen only what it consistently gets right, and keep anything touching a commitment, a payoff figure or a disbursement with a licensed person.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
How an industrial equipment OEM builds a parts-desk test set from its own closed orders, grades the agent on as-built revisions, and stops wrong-part shipments.
Community management agents fail on escalations and claims, not tone. Build a 400-thread test set from your own social inbox before anything goes live.
Cattle producers can now measure a phone agent against real past calls, watch every step it took, and widen what it is allowed to do only once it earns it.
Sort six months of front-counter recordings into eight call types, write the right answer for each, then widen the agent's authority in four stages, not one.
How a design firm builds an RFI test set from closed jobs, scores citation and routing accuracy, and widens an agent's authority only when the numbers earn it.
Six tests that make a title order routine, the escalation list that never bends, who owns the rule, and a costed 90-file month showing where the money sits.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI