By Sagar Shankaran, Founder of CallSphere
Build an AI test set from closed mortgage files: the 1008 qualifying income is the answer key, QC defects are the hard cases, and the breakdown beats the score.
Key takeaways
Plenty of mortgage shops bought a document-reading tool two years ago, watched it pull the wrong figure off a commission-heavy paystub, and quietly stopped using it. The tool was not the whole problem. The problem was that nobody could tell you, before it touched a live file, how often it would be wrong or on which kinds of borrowers. You bought it, turned it on, and found out from your post-closing quality control report.
That is the thing that actually changed in 2026. The tooling for watching an agent work and scoring it against real cases matured to the point where it now gates the rollout rather than following it. You measure the agent against files where you already know the right answer, you read back every step it took to get there, and only then do you widen what it is allowed to do. It is the same discipline your underwriting manager applies to a new underwriter: they audit the first fifty files at 100%, not because they doubt the person, but because that is how you find out what they do not know yet.
Here is the plain version: before an agent touches a borrower's file, you replay it against a few hundred loans you already closed, compare its answer to the answer your underwriter wrote on the 1008, and count the misses. That set of loans is your test set, and you already own it.
Every closed loan in your shop has a Uniform Underwriting and Transmittal Summary — the 1008 — with the final qualifying income your underwriter used. Not the income the borrower claimed, not the number on the 1003. The number a licensed human signed off on after reading the paystubs, the W-2s, the tax transcripts from the 4506-C, the K-1s, and whatever letter of explanation was needed. That is an answer key you did not have to build. It cost you nothing and it is specific to your credit box, your overlays, and the kind of borrower your loan officers actually bring in.
The second half of the key is your quality control file. Fannie Mae's selling guide has you reviewing a random sample of closed loans every month plus a discretionary sample. Income miscalculation is one of the most common defects that turns up in those reviews. Those defect files are gold for this purpose. They are the exact cases where a competent human already got it wrong once. If your agent gets those right, that tells you something. If it gets them wrong the same way, that tells you more.
Do not build the test set out of clean files. A test set of 200 salaried W-2 borrowers with two-year job histories proves nothing except that easy is easy.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for financial services in your browser — 60 seconds, no signup.
Pull from the last twelve months so seasonal mix is represented — the spring purchase run, the fall slowdown, whatever refinance burst happened when rates moved. Aim for a set that looks roughly like this: 90 salaried borrowers, 45 self-employed on Schedule C or with S-corp and partnership K-1s, 25 with rental income on Schedule E, 20 with variable pay (commission, bonus, overtime, shift differential), 12 government loans with their own quirks, and every single file that produced a QC defect or a repurchase inquiry, even if that pushes you past 200.
flowchart TD
A["Pull 200 closed loans from the last 12 months"] --> B["Lock the 1008 qualifying income as the answer key"]
B --> C["Replay the agent file by file, no hints"]
C --> D{"Within $50 a month of the underwriter?"}
D -->|Misses on 9 or more files| E["Read the step list, fix the rule, log what it missed"]
E --> C
D -->|191 or better| F["Shadow mode on live files for two weeks"]
F --> G["Let it write to the conditions log on one doc type"]
The loop back to the replay step is the part shops skip. You do not fix the agent once. You fix it, run the whole set again, and see whether the fix broke something that was working. Ask any vendor whether their tool lets you re-run the same set on demand. If the answer is a demo instead of a yes, you have your answer.
Scoring alone will fool you. An agent can produce the right monthly income on a rental property file by adding two numbers that happen to sum correctly while completely ignoring the vacancy factor and the fact that the property was acquired in June. Right answer, wrong reason, and it will fail the moment the next file has a full year of rents.
The 2026 tooling shows you the work: which document it opened, which page it read the figure from, what it did with the depreciation add-back, where it decided a bonus was not stable enough to count. You read that trail on the files it got right as well as the ones it missed. Budget an underwriter for two afternoons on this — it is the highest-value two afternoons anyone in your shop will spend on AI this year.
Watch for one specific failure that shows up in this trade: the agent averaging twenty-four months of self-employed income when the trend is declining. The arithmetic is right, the guideline treatment is wrong, and a declining trend is precisely the thing your underwriter is paid to notice.
Illustration with stated assumptions. Say the agent lands within $50 a month of the 1008 on 192 of 200 files — a 4% miss rate — and your shop closes 150 loans a month.
| Line | Assumption | Result |
|---|---|---|
| Test files | Stated | 200 |
| Within $50/month of the 1008 | 96% | 192 |
| Misses | 4% | 8 files |
| Misses on salaried W-2 borrowers | 1 of 90 | 1.1% |
| Misses on self-employed borrowers | 6 of 45 | 13.3% |
| Misses on rental income files | 1 of 25 | 4.0% |
| Monthly closings | Stated | 150 |
| Salaried share of closings | 60% | 90 files |
| Expected misses if used on salaried only | 1.1% of 90 | about 1 file/month |
The headline number, 4%, is the useless one. The breakdown is the decision. A 1.1% miss rate on salaried borrowers, caught by a processor who still verifies, is workable. A 13.3% miss rate on self-employed income means the agent does not touch self-employed files this quarter, full stop. Same tool, two different answers, and you only get to that conclusion because you sliced the test set by borrower type before you turned anything on.
It will not tell you whether the borrower is lying. A fabricated paystub that matches a fabricated Work Number record will pass every consistency check the agent runs, because the agent's whole job is checking documents against each other. Fraud detection is a separate discipline and it stays with your underwriter and your vendor for verification.
Still reading? Stop comparing — try CallSphere live.
See the financial services AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
It will not tell you anything about fair lending. Two hundred closed loans is a sample of the people you approved. It contains no information about the applicants you declined or the ones who never applied. Anyone who tells you a test set of funded loans proves an agent is unbiased is selling something. Fair lending review is a compliance function, it uses different data, and it belongs to your compliance officer.
And it will not survive a guideline change. When an agency updates its treatment of a particular income type, or you add an overlay, your test set answers are now partly stale. Re-run it after every material guideline change and after any vendor tells you the model was updated. Put that on the calendar the same way you put the annual policy review on it.
Start with forty and weight them toward the hard types — twenty self-employed, ten rental, ten variable pay. Forty files will not give you a defensible miss rate, but it will tell you within a day whether the tool is anywhere near ready. If it misses eleven of forty, you just saved yourself a two-hundred-file exercise.
Your underwriting manager runs it and a senior processor does the file pulling. This is not a technical job. It is reading a number the agent produced, comparing it to a number on a form, and writing down which one is right. If a vendor's tool requires someone technical to score it, that is a mark against the tool.
Yes, and that is a real consideration. These are closed files with full personal information in them. Handle the test the way you handle any other access to the loan file: named users, logged access, the vendor under your existing information-security agreement, and no copies of tax transcripts sitting in someone's downloads folder. Your information security policy already covers this — apply it rather than writing a new one.
Quarterly as a baseline, plus immediately after any guideline change, any overlay change, and any notice from the vendor that the model behind it changed. Add ten fresh files from the current quarter each time so the set does not go stale and start rewarding an agent that has effectively memorized your old cases.
Ask post-closing for forty funded loans from the last six months, weighted toward self-employed and rental income, with the 1008 in each one. That is your first afternoon. Everything else — the shadow period, the conditions log, the widening — follows from what those forty files tell you, and none of it requires you to trust a demo.
The same standard applies to anything you put on the phone. If you are considering an AI agent for the front line, ask to hear it against your own recorded calls before it answers a live borrower. CallSphere builds AI voice and chat agents that answer the branch line and web chat after hours, book the appointment with the loan officer, and capture the lead — and every call is transcribed, so you can score it the same way you would score a new hire's first fifty calls.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
The 200-call test set a treatment program should build from its own recordings, the five things to score, and how to widen an agent's authority safely.
Freezer logs, eyewash tags, reagent expiry and MRI Zone III: what autonomous inspection rounds actually cover in a diagnostic building, and the honest math.
An AI agent in a lending shop should read widely, write to the conditions log, and send nothing with a routing number. The permissions to remove this Monday.
Build a 200-case test set from your own janitorial work orders, grade the agent, and widen its authority in stages. Includes the five pass/fail call types.
Community bank hiring changed in 2026: memorizing the fee schedule is out, verifying a cited answer against the core screen is in. A week-one plan and the ramp math.
How a caterer builds a 240-case test set from real BEOs and inquiry email, what it costs to grade it, and the category floors to hit before going live.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI