By Sagar Shankaran, Founder of CallSphere
SAR confidentiality rules out cloud chatbots. Here is what an on-site model does for a credit union BSA officer, the hours math, and the examiner questions.
Key takeaways
It is 4:15 on a Tuesday afternoon in the second-floor operations room. Your BSA officer has a Verafin alert open on one screen, ninety days of a member's transaction history pulled from the core on the other, and a branch referral form — the one the Elm Street teller filled out by hand because the member kept coming in with $9,400 in cash on Fridays — sitting next to her keyboard. She has been at this narrative for fifty minutes. It is the eleventh one this quarter, the detection date was nineteen days ago, and the filing clock does not care that she is also the compliance officer, the OFAC administrator and the person who runs annual staff training.
Every other department in your credit union has spent 2026 pasting work into a chatbot to get a first draft. She cannot. And she is right not to.
A Suspicious Activity Report and its supporting material are confidential by federal law. You may not disclose a SAR, or even the fact that one was filed, to anyone outside the narrow set of people entitled to know. Pasting the narrative — or the case facts assembled to write it — into a consumer chat product operated by a third party is not a gray area your compliance committee gets to reason its way through. Layer on the GLBA Safeguards obligations that NCUA enforces through Part 748 and its appendices, your member privacy notice, and the vendor due diligence file your examiner asks for by name, and the honest answer for the last two years was: the BSA desk does this by hand.
On-premises AI means the model runs on a machine inside your building, on your network, so the member data it reads never travels to anyone else's servers — the file goes to the model instead of the model's owner getting the file. That is the sentence that opens this desk back up, and 2026 is the year it became practical at credit union scale rather than bank-holding-company scale.
Two things. First, the chips. Qualcomm's Dragonwing-class processors and their peers made local processing genuinely capable on ordinary equipment, so "runs on site" no longer means a refrigerated room and a six-figure purchase order. Second, and more persuasive to a board: large enterprises stopped treating on-premises as the legacy option. Cisco is rolling a personal AI agent out to roughly 90,000 employees with explicit emphasis on on-premises operation for control and data protection. When a company that size chooses to keep the processing inside its own walls, your examiner is not going to look at your $9,000 server and ask why you did not use the cloud.
Combine that with the open models discussed everywhere this year, and the practical result is that a credit union under a billion in assets can now put a capable model on a box in the ops room, wire it to a folder that only the BSA team can reach, and let it read case files that legally cannot leave the premises.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent for financial services in your browser — 60 seconds, no signup.
flowchart TD
A["Monitoring alert lands in the BSA queue"] --> B["Analyst pulls 90 days of activity from the core"]
A --> C["Branch referral form scanned into the case folder"]
B --> D["On-site model drafts the who-what-when-where-why narrative"]
C --> D
D --> E["BSA Officer edits, verifies every date and dollar amount"]
E --> F["Filed through FinCEN BSA E-Filing inside the deadline"]
F --> G["Case file and draft stay on the in-house server"]
It does not decide anything. It reads. Give it the transaction extract, the referral form, the prior case notes and the member's account opening record, and ask it to produce a chronological narrative in the structure FinCEN expects: who conducted the activity, what instruments and amounts, when it occurred and over what period, where it happened — which branch, which ATM, which shared branching location — and why your institution considers it suspicious. Your officer gets a draft with the dates and dollar figures already assembled in order, which is the part that eats her afternoon.
It is also useful on the boring end. A cash-intensive member — a car wash, a laundromat, a used car lot — generates dozens of Currency Transaction Reports a year and a steady drip of alerts that turn out to be exactly what you would expect from a car wash. Reading three months of activity and writing four sentences explaining why an alert was cleared, so that your annual independent testing has something to look at, is precisely the work that gets deferred until the auditor shows up in September.
Assumptions, illustrative: 19,000 members, 34 SARs filed a year, roughly 220 alerts triaged a year, a BSA officer at a loaded $61 an hour, and one part-time analyst.
| Task | By hand today | With an on-site draft | Yearly hours |
|---|---|---|---|
| Narrative writing, 34 SARs | 2.5 hrs each | 1.1 hrs each | 85 → 37 |
| Alert clearing notes, 220 alerts | 25 min each | 11 min each | 92 → 40 |
| Total | 177 hrs | 77 hrs | 100 hrs saved |
At $61 an hour that is about $6,100 a year of officer time. A server sized for this work, amortized over three years, plus power, runs somewhere around $5,000 to $6,000 a year at the low end. So on BSA alone, this roughly breaks even — and any vendor telling you the labor savings pay for the hardware is selling you something.
The case is elsewhere. First, on the deadline. A late SAR is a documented BSA program weakness that goes into your examination report and follows you for years; the hundred hours you free up are the hours that keep you off the wrong side of the clock in a quarter when your officer is also out for two weeks. Second, on the box itself: it is not a BSA box. Once it is in the room, collections uses it for hardship letters, HR uses it for job postings, and lending uses it for policy questions. Spread across four departments, the arithmetic stops being marginal and starts being obvious.
Have this ready before you turn the machine on. Who has access to it, and how is that access removed at termination? What is written in your board-approved BSA program about the use of drafting tools? Who reviews and signs the final narrative — a person, named? Is there a record showing the officer edited and approved each filing rather than submitting a draft as written? Where does the draft live, how long is it retained, and how does that square with your record retention schedule under Part 749?
Answer those five, in writing, and this is a short conversation. Skip them and you have turned a productivity gain into an examination finding, which is a spectacularly bad trade.
Still reading? Stop comparing — try CallSphere live.
See the financial services AI agent handle a real call — complete, industry-specific, and live in your browser. No signup.
The decision to file. That belongs to a human being with a name, a title and accountability, and it should stay that way even if the drafting is assisted. Do not let a tool decide an alert is unremarkable and close it unread — false negatives are the whole risk in BSA, and they are invisible until an examiner or a subpoena finds them.
Also keep humans on the sensitive judgment calls: suspected elder financial exploitation, where the right next step is often a phone call to the member's branch manager and possibly Adult Protective Services rather than a filing; anything involving an employee; and 314(b) information sharing with another institution. And check every number. A drafting tool that assembles dates and amounts will occasionally transpose one, and a SAR with the wrong dollar figure is worse than a late one.
Do not start with a purchase. Start with a folder. Take three closed cases from last year — already filed, already resolved — and have your officer time herself rewriting the narratives with and without a draft. Two afternoons. If the draft does not cut her time roughly in half on real cases, the hardware conversation is premature. If it does, take the timing sheet to your CUSO or data processing partner and ask what it would cost them to run a model for you on equipment that sits in your building or theirs under your control.
You can for most departments. For SAR content specifically, the calculation is different, because you are not weighing convenience against privacy policy — you are weighing it against a federal confidentiality rule and an examiner who will ask where the file went. If your general counsel or compliance consultant signs off in writing on a specific cloud arrangement, fine. Absent that sign-off, keeping it in the building is the defensible answer.
No. From her seat it is a text box on the internal network. The technical ownership sits with whoever administers your servers, and the practical question you should be asking is who patches it and who restores it. If that person is your only core administrator, get the hosting done by your data processing partner instead.
Your monitoring system decides what to look at — it scores transactions and raises alerts. What is described here writes about what you have already decided to look at. They are complementary, and you should not replace tuned monitoring rules with a general model. Keep the alerting where it is.
Indirectly and meaningfully. The most common finding in a BSA independent test at a shop your size is thin documentation on cleared alerts — not that the judgment was wrong, but that nobody wrote down why. Cutting the cost of writing that paragraph from twenty-five minutes to eleven is what actually gets it written.
One adjacent point worth raising with your operations team: BSA and fraud work generate outbound and inbound member calls at exactly the wrong hours — the member whose card was blocked, the member calling back about a hold. CallSphere builds AI voice and chat agents that answer the member line around the clock, handle the routine questions, book a callback with the right person, and capture the details before a human picks up. It does not touch case files or filings; it keeps the phones from burying the people who are working them.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Supplement brands get adverse events reported by ticket, not by form. On-premises AI reads all 3,400 a month without that text ever leaving the building.
Concealed-damage claims cost a warehouse manager half a Friday. On-site AI searches your own footage and scan trail without any data leaving the property.
Casino surveillance footage cannot leave the property. On-premises AI in 2026 cuts a 90-minute disputed handpay review to nine, with no video sent to a vendor.
Open models closed the gap in 2026. A worked buy-versus-rent comparison for a 58-employee credit union, including the staffing cost nobody budgets for.
Custody orders, health plans, subsidy files and classroom video can stay in the building. What on-premises AI unlocks for a child care center in 2026.
A credit union's core agreement plus nine amendments now fits in one question. Find the notice window, the escalators and the exit fees before renewal.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.
Try Live DemoBook a DemoCalculate Your ROI