By Sagar Shankaran, Founder of CallSphere
How a design firm builds an RFI test set from closed jobs, scores citation and routing accuracy, and widens an agent's authority only when the numbers earn it.
Key takeaways
Somebody pasted an RFI into a chat box during the parking garage job. The answer came back in nine seconds, beautifully written, citing a detail number that does not exist on the drawings. The project architect spent longer checking it than she would have spent answering the RFI herself, and the firm went back to the old way. If that is your experience, you are not wrong and you are not behind.
What changed in 2026 is not the writing. It is that you can measure an agent against your own closed files before it touches a contractor, watch every step it took, and widen what it may do only when the numbers earn it. That step is the difference between a party trick and something you would let near construction administration.
Construction administration is where A/E fee goes to die. You priced CA as a percentage back when the job was a schematic idea, and now the general contractor issues eleven RFIs a week on a project you budgeted for four. Your agreement — AIA B101 tied to A201, or the owner's version of it — gives you a reasonable time to respond, and specification section 01 31 00 usually pins that to seven or ten calendar days. Miss it and the contractor's scheduler starts building a delay narrative in the monthly update, copied to the owner.
The real cost is not the delay claim. It is that a project architect billing at $165 earns nothing extra for the RFI that should never have existed — the one asking what the drawings already answer, or asking for a redesign for free under the word "clarification." A firm with 350 RFIs on a 12-month CA phase burns 300 to 500 hours on work priced at half that.
Agent evaluation, in plain terms, is grading a piece of software against a pile of decisions you already made and know the outcome of, before it makes a single new one. In construction administration, you have an unusually good pile: every RFI you closed last year, with the response you actually gave, the drawing and spec section you cited, and whether it turned into an ASI, a proposal request, or nothing at all.
Through 2025, the tools for watching an agent work were built for engineers who write software. In 2026 that measurement became the ordinary gate on a rollout — what you do before widening authority, not after something goes wrong. Three consequences for a design firm.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent in your browser — 60 seconds, no signup.
First, you can run a batch: feed it 300 old RFIs it has never seen, answers withheld, and get a scorecard against what the architect of record actually wrote. Second, you can see the work, not just the answer — which sheet it opened, which spec section it read, which prior RFI it matched. When it gets one wrong you can tell whether it misread the detail or whether your own set was genuinely ambiguous, and that second case is worth knowing regardless. Third, authority turns up in steps: draft-only, then draft-and-send for one category, then more.
None of that requires a developer. It requires somebody in your office — usually the CA lead or a senior project architect — willing to spend two days building a test set out of the firm's own history.
flowchart TD
A["Export 300 closed RFIs from Procore or Newforma"] --> B["Label each: clarification, spec conflict, substitution, design change"]
B --> C["Run the agent blind, answers withheld"]
C --> D{"Did it cite the same sheet and spec section the architect cited?"}
D -->|Below 90 percent| E["Fix the source set: superseded sheets, missing addenda, unindexed ASIs"]
E --> C
D -->|90 percent or better| F["Draft-only mode on a live job, project architect signs every response"]
F --> G{"Four weeks with no wrong citation?"}
G -->|No| E
G -->|Yes| H["Agent sends routine clarifications, PM reviews the daily log"]
Do not grab the most recent 300. Build the set deliberately — the point is to find where it fails, not to get a good grade. From your last four or five closed jobs, take:
Feed it what a project architect would have had: the conformed set, the addenda, every ASI issued to date, the approved submittals, the specification. If your firm cannot assemble that quickly for a closed job, that is finding number one, and it has nothing to do with AI.
Grade on three things and nothing else. Correct citation: same sheet, detail and spec section the architect of record pointed at? Correct routing: did it spot a design change, substitution or field condition and refuse to answer it as a clarification? No invention: did it name any detail, section or product that does not exist in the set?
Assumptions for the arithmetic: 300 test RFIs, a CA phase that generates 350 RFIs a year across the office, 25 minutes of project architect time per RFI at a burdened $118 an hour, and a target that the agent handles only the clarification category, which is 40% of volume.
| Measure | First run | After fixing the document set | Gate to pass |
|---|---|---|---|
| Correct citation, clarifications | 81% | 94% | 90% |
| Correct routing of design changes | 62% | 88% | 95% |
| Invented a detail or product | 7 of 300 | 0 of 300 | 0 |
| RFIs it may handle | none | clarifications only, draft | — |
Read that table properly. Citation accuracy passed. Routing did not — 88% means 12 design changes in 100 would have been answered as free clarifications, which is exactly how a firm gives away redesign work. So the agent goes to draft-only on clarifications, every draft signed by a human, and you re-run the routing test in a month. On labor, 140 clarification RFIs a year at 25 minutes each is 58 hours; halving the drafting portion is about 29 hours, roughly $3,400. Not a headline number — a real one, and it grows as the categories you trust widen.
Anything touching life safety stays with the licensed professional. Egress width, fire-rated assembly substitutions, structural member changes, guardrail geometry — the agent may retrieve the sheet and the code section, then it stops. Anything that changes contract sum or time stays human, because that is a commercial decision your project manager makes with the owner.
Still reading? Stop comparing — try CallSphere live.
CallSphere ships complete AI voice agents per industry — 14 tools for healthcare, 10 agents for real estate, 4 specialists for salons. See how it actually handles a call before you book a demo.
Anything with a claim smell stays human. When a contractor's RFI is really a letter — recounting a sequence of events, copied to the owner — that is not a question, it is the opening of a dispute record, and your liability carrier wants a person answering it. The rule for your team: more narrative than question, it goes to the project manager.
And the seal never moves. An agent does not stamp drawings, issue an ASI, or sign a certificate of substantial completion. Your state board licenses people, not software, and the responsible charge rules have no exception for a good scorecard.
A hundred shows you the shape of the failures; 300 lets you trust a percentage. With only 80 closed RFIs, use all of them and be honest that the score has wide error bars. Coverage matters more than count — make sure design changes dressed up as clarifications are in there.
No, but it costs a step. Most firms export the RFI log to a spreadsheet at closeout anyway for the project record. Pair that export with the conformed set and your own ASI log and you have a test set. If the contractor's system is the only record of the response you gave, fix that for reasons unrelated to any of this.
The truth, which for the first year is simple: responses are drafted with software assistance and reviewed and issued by the project architect. Nothing in the contract changes, because the responsible party has not. Some owner agreements — federal and healthcare system contracts especially — now carry AI use clauses, so read the ones you signed in the last eighteen months.
Later than you think, and possibly never outside the narrowest category. A defensible sequence: draft-only for a full CA phase, then auto-send for pure "see detail X, no change to contract sum or time" responses on one job, with the project manager reading the daily log. Re-run the routing test monthly — a new contractor changes the mix of what walks in the door.
Pick one closed job. Export its RFI log with responses to a spreadsheet and have someone label thirty by category — clarification, conflict, field condition, substitution, design change. That is the whole first session, and it will already tell you something uncomfortable about how many of last year's "clarifications" were free design work. Then run those thirty past the agent with the answers hidden. If the routing is bad, you know what to fix before anything goes live; if it is good, build the set out to 300.
Every RFI has a phone call attached to it — a superintendent asking whether it was answered, a supplier chasing a submittal, an owner's rep calling about a pay application — and those calls hit a small firm's line at the worst hours. CallSphere builds AI voice and chat agents that answer the firm's phone and web chat around the clock, capture exactly which project and RFI number the caller means, book the call-back on the right project architect's calendar, and hand over a transcript. It does not answer the RFI — a licensed professional does — but it stops the answer being delayed by voicemail tag.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
How an industrial equipment OEM builds a parts-desk test set from its own closed orders, grades the agent on as-built revisions, and stops wrong-part shipments.
Community management agents fail on escalations and claims, not tone. Build a 400-thread test set from your own social inbox before anything goes live.
Cattle producers can now measure a phone agent against real past calls, watch every step it took, and widen what it is allowed to do only once it earns it.
Sort six months of front-counter recordings into eight call types, write the right answer for each, then widen the agent's authority in four stages, not one.
Build a 300-call test set from your own cancellations and answering-service log, grade escalation at 100%, then widen the agent one visit type at a time.
How A/E firms hand SF-330 assembly to Claude Cowork or ChatGPT Work: what it drafts, the worked hours math, and the profiles a licensed human must still read.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI