---
title: "Before an AI Agent Touches a Condition, Replay It Against 200 Closed Files Where You Know the Qualifying Income"
description: "Build an AI test set from closed mortgage files: the 1008 qualifying income is the answer key, QC defects are the hard cases, and the breakdown beats the score."
canonical: https://callsphere.ai/blog/before-an-ai-agent-touches-a-condition-replay-it-against-200-closed-fi
category: "Financial Services"
tags: ["mortgage lending", "ai evaluation", "quality control", "income calculation", "underwriting", "loan operations"]
author: "CallSphere Team"
published: 2026-06-28T08:07:13.000Z
updated: 2026-09-06T02:57:35.235Z
---

# Before an AI Agent Touches a Condition, Replay It Against 200 Closed Files Where You Know the Qualifying Income

> Build an AI test set from closed mortgage files: the 1008 qualifying income is the answer key, QC defects are the hard cases, and the breakdown beats the score.

## You Tried a Document Reader in 2024 and It Invented a Pay Rate. That Objection Is Fair.

Plenty of mortgage shops bought a document-reading tool two years ago, watched it pull the wrong figure off a commission-heavy paystub, and quietly stopped using it. The tool was not the whole problem. The problem was that nobody could tell you, before it touched a live file, how often it would be wrong or on which kinds of borrowers. You bought it, turned it on, and found out from your post-closing quality control report.

That is the thing that actually changed in 2026. The tooling for watching an agent work and scoring it against real cases matured to the point where it now gates the rollout rather than following it. You measure the agent against files where you already know the right answer, you read back every step it took to get there, and only then do you widen what it is allowed to do. It is the same discipline your underwriting manager applies to a new underwriter: they audit the first fifty files at 100%, not because they doubt the person, but because that is how you find out what they do not know yet.

Here is the plain version: **before an agent touches a borrower's file, you replay it against a few hundred loans you already closed, compare its answer to the answer your underwriter wrote on the 1008, and count the misses.** That set of loans is your test set, and you already own it.

## The Answer Key Is Sitting in Your Post-Closing Files

Every closed loan in your shop has a Uniform Underwriting and Transmittal Summary — the 1008 — with the final qualifying income your underwriter used. Not the income the borrower claimed, not the number on the 1003. The number a licensed human signed off on after reading the paystubs, the W-2s, the tax transcripts from the 4506-C, the K-1s, and whatever letter of explanation was needed. That is an answer key you did not have to build. It cost you nothing and it is specific to your credit box, your overlays, and the kind of borrower your loan officers actually bring in.

The second half of the key is your quality control file. Fannie Mae's selling guide has you reviewing a random sample of closed loans every month plus a discretionary sample. Income miscalculation is one of the most common defects that turns up in those reviews. Those defect files are gold for this purpose. They are the exact cases where a competent human already got it wrong once. If your agent gets those right, that tells you something. If it gets them wrong the same way, that tells you more.

Do not build the test set out of clean files. A test set of 200 salaried W-2 borrowers with two-year job histories proves nothing except that easy is easy.

## What Goes in the 200 Files

Pull from the last twelve months so seasonal mix is represented — the spring purchase run, the fall slowdown, whatever refinance burst happened when rates moved. Aim for a set that looks roughly like this: 90 salaried borrowers, 45 self-employed on Schedule C or with S-corp and partnership K-1s, 25 with rental income on Schedule E, 20 with variable pay (commission, bonus, overtime, shift differential), 12 government loans with their own quirks, and every single file that produced a QC defect or a repurchase inquiry, even if that pushes you past 200.

```mermaid
flowchart TD
  A["Pull 200 closed loans from the last 12 months"] --> B["Lock the 1008 qualifying income as the answer key"]
  B --> C["Replay the agent file by file, no hints"]
  C --> D{"Within $50 a month of the underwriter?"}
  D -->|Misses on 9 or more files| E["Read the step list, fix the rule, log what it missed"]
  E --> C
  D -->|191 or better| F["Shadow mode on live files for two weeks"]
  F --> G["Let it write to the conditions log on one doc type"]
```

The loop back to the replay step is the part shops skip. You do not fix the agent once. You fix it, run the whole set again, and see whether the fix broke something that was working. Ask any vendor whether their tool lets you re-run the same set on demand. If the answer is a demo instead of a yes, you have your answer.

## Reading the Step List, Not Just the Score

Scoring alone will fool you. An agent can produce the right monthly income on a rental property file by adding two numbers that happen to sum correctly while completely ignoring the vacancy factor and the fact that the property was acquired in June. Right answer, wrong reason, and it will fail the moment the next file has a full year of rents.

The 2026 tooling shows you the work: which document it opened, which page it read the figure from, what it did with the depreciation add-back, where it decided a bonus was not stable enough to count. You read that trail on the files it got right as well as the ones it missed. Budget an underwriter for two afternoons on this — it is the highest-value two afternoons anyone in your shop will spend on AI this year.

Watch for one specific failure that shows up in this trade: the agent averaging twenty-four months of self-employed income when the trend is declining. The arithmetic is right, the guideline treatment is wrong, and a declining trend is precisely the thing your underwriter is paid to notice.

## Scoring It: What a 4% Miss Rate Actually Costs You

Illustration with stated assumptions. Say the agent lands within $50 a month of the 1008 on 192 of 200 files — a 4% miss rate — and your shop closes 150 loans a month.

| Line | Assumption | Result |
| --- | --- | --- |
| Test files | Stated | 200 |
| Within $50/month of the 1008 | 96% | 192 |
| Misses | 4% | 8 files |
| Misses on salaried W-2 borrowers | 1 of 90 | 1.1% |
| Misses on self-employed borrowers | 6 of 45 | 13.3% |
| Misses on rental income files | 1 of 25 | 4.0% |
| Monthly closings | Stated | 150 |
| Salaried share of closings | 60% | 90 files |
| Expected misses if used on salaried only | 1.1% of 90 | about 1 file/month |

The headline number, 4%, is the useless one. The breakdown is the decision. A 1.1% miss rate on salaried borrowers, caught by a processor who still verifies, is workable. A 13.3% miss rate on self-employed income means the agent does not touch self-employed files this quarter, full stop. Same tool, two different answers, and you only get to that conclusion because you sliced the test set by borrower type before you turned anything on.

## What the Test Set Will Never Tell You

It will not tell you whether the borrower is lying. A fabricated paystub that matches a fabricated Work Number record will pass every consistency check the agent runs, because the agent's whole job is checking documents against each other. Fraud detection is a separate discipline and it stays with your underwriter and your vendor for verification.

It will not tell you anything about fair lending. Two hundred closed loans is a sample of the people you approved. It contains no information about the applicants you declined or the ones who never applied. Anyone who tells you a test set of funded loans proves an agent is unbiased is selling something. Fair lending review is a compliance function, it uses different data, and it belongs to your compliance officer.

And it will not survive a guideline change. When an agency updates its treatment of a particular income type, or you add an overlay, your test set answers are now partly stale. Re-run it after every material guideline change and after any vendor tells you the model was updated. Put that on the calendar the same way you put the annual policy review on it.

## Frequently asked questions

### Two hundred files sounds like a lot of work. Can I start smaller?

Start with forty and weight them toward the hard types — twenty self-employed, ten rental, ten variable pay. Forty files will not give you a defensible miss rate, but it will tell you within a day whether the tool is anywhere near ready. If it misses eleven of forty, you just saved yourself a two-hundred-file exercise.

### Who in my shop should run this? I don't have anyone technical.

Your underwriting manager runs it and a senior processor does the file pulling. This is not a technical job. It is reading a number the agent produced, comparing it to a number on a form, and writing down which one is right. If a vendor's tool requires someone technical to score it, that is a mark against the tool.

### Does the agent need access to real borrower documents to be tested?

Yes, and that is a real consideration. These are closed files with full personal information in them. Handle the test the way you handle any other access to the loan file: named users, logged access, the vendor under your existing information-security agreement, and no copies of tax transcripts sitting in someone's downloads folder. Your information security policy already covers this — apply it rather than writing a new one.

### How often do I re-run the test after go-live?

Quarterly as a baseline, plus immediately after any guideline change, any overlay change, and any notice from the vendor that the model behind it changed. Add ten fresh files from the current quarter each time so the set does not go stale and start rewarding an agent that has effectively memorized your old cases.

## Pull Forty Files This Week

Ask post-closing for forty funded loans from the last six months, weighted toward self-employed and rental income, with the 1008 in each one. That is your first afternoon. Everything else — the shadow period, the conditions log, the widening — follows from what those forty files tell you, and none of it requires you to trust a demo.

The same standard applies to anything you put on the phone. If you are considering an AI agent for the front line, ask to hear it against your own recorded calls before it answers a live borrower. [CallSphere](https://callsphere.ai) builds AI voice and chat agents that answer the branch line and web chat after hours, book the appointment with the loan officer, and capture the lead — and every call is transcribed, so you can score it the same way you would score a new hire's first fifty calls.

---

Source: https://callsphere.ai/blog/before-an-ai-agent-touches-a-condition-replay-it-against-200-closed-fi
