---
title: "AI Safety Institute Pre-Launch Evaluation Framework: What It Covers"
description: "The US AI Safety Institute's pre-launch evaluation framework — what it tests, what it means for enterprise buyers, and how CallSphere maps to it."
canonical: https://callsphere.ai/blog/tw26w19-ai-safety-institute-pre-launch-evaluation-framework
category: "Guides & News"
tags: ["AI Safety Institute", "AI Policy", "CAISI", "AI Evaluation", "Enterprise AI"]
author: "CallSphere Team"
published: 2026-05-04T00:00:00.000Z
updated: 2026-05-18T01:11:29.662Z
---

# AI Safety Institute Pre-Launch Evaluation Framework: What It Covers

> The US AI Safety Institute's pre-launch evaluation framework — what it tests, what it means for enterprise buyers, and how CallSphere maps to it.

## A New Gatekeeper for Frontier Models

This week the US AI Safety Institute (now operating under the post-CAISI compliance umbrella) formalized a **pre-launch evaluation framework** — third-party testing of frontier models before public release. Major labs (OpenAI, Anthropic, Google DeepMind, Meta) have agreed to share pre-release weights and capabilities information with the Institute for evaluation.

For enterprise buyers, this is the first time there is a clear, public-sector signal about whether a model has been independently stress-tested before it shows up in your stack.

## What the Framework Tests

The framework groups evaluations into four buckets:

- **Capability evaluations** — what can the model actually do, across reasoning, coding, agentic tasks, multilingual performance, and tool use
- **Misuse evaluations** — biothreat uplift, cyber-offensive uplift, CBRN, persuasion / manipulation
- **Autonomy & control** — self-replication, resource acquisition, deception under pressure
- **Bias and fairness** — disparate-impact testing across protected classes

Each category produces a structured report shared with the lab and (in summary form) with the public. Frontier labs cannot launch a covered model without completing the cycle.

## Why Buyers Should Care

You have probably bought a SaaS product that uses GPT-class or Claude-class models. Three things shift once pre-launch eval is the norm:

1. **Lower tail risk on net-new releases.** A model that passed an Institute eval is less likely to surprise you a month after rollout.
2. **Vendor due diligence gets easier.** You can ask "which model are you running, and did its base pass an AISI eval?" and get a real answer.
3. **Audit trails strengthen.** When regulators ask why you trusted an AI vendor in 2026, "the underlying model went through pre-launch evaluation" is now part of the answer.

## What It Does Not Cover

The framework evaluates **base models**, not every product built on top of them. Your AI voice agent vendor still has to do their own evaluation work — prompt safety, tool-use sandboxing, data handling, retention, vertical-specific risk (PHI, PCI, etc.).

This is where many enterprise programs get tripped up: they assume the base-model eval flows downstream automatically. It does not.

## The Stack View

```mermaid
flowchart TB
    AISI[AI Safety Institute
pre-launch eval] --> Base[Base model
GPT, Claude, Gemini]
    Base --> Plat[AI platform layer
your vendor]
    Plat --> App[Your deployment]
    App --> User[End user]
    AppEval[Vendor-side evals
prompt safety, tool sandbox, vertical risk] --> App
```

The Institute owns the top. Your vendor owns the middle. You own deployment. None of those three layers can substitute for the other.

## How CallSphere Maps to This

CallSphere runs on third-party frontier models (OpenAI, Anthropic) plus our own orchestration layer. What that means in practice:

- Base-model risk is reduced by the Institute's framework before models ever reach our platform
- We add vertical-specific guardrails on top: HIPAA-friendly handling for healthcare, after-hours / sales / IT helpdesk / salon / real estate prompt templates
- 14 function tools are sandboxed; voice/chat/SMS/WhatsApp channels each have their own retention and consent defaults
- 57+ languages share the same guardrail stack
- 3–5 day launch keeps the human-in-the-loop tight — you see what the agent says before it goes live

You inherit the pre-launch eval at the base layer. We do the rest.

## What to Ask Your AI Vendor This Quarter

A short checklist to bring to your next vendor review:

- Which base model(s) do you run, and have those models been through an AISI pre-launch eval?
- Where is your platform-layer evaluation documented?
- What is your refresh cycle when a base model updates?
- Do you support vertical-specific prompt and tool restrictions?
- What is your incident-response timeline if a model regression hits production?

If you cannot get clean answers to those five questions, the vendor is not enterprise-ready in 2026.

## The Real Win

The Institute framework is not the end of AI safety, but it is the first thing that gives enterprise buyers a defensible "we did our homework" answer for boards and regulators. Use it.

## CTA

If you are evaluating an AI voice or chat platform and need a vendor that takes pre-launch model evaluation, vertical guardrails, and HIPAA-friendly deployment seriously — start a free trial at [https://callsphere.ai/trial](https://callsphere.ai/trial) or book a demo.

## FAQ

**Q: Does an AISI eval mean a model is "safe"?**
A: No. It means the model has been independently stress-tested on a defined set of capabilities and risks. Safety in production still depends on how the model is deployed and constrained.

**Q: Will the framework apply to non-US labs?**
A: Coverage today is voluntary and US-centric, but the major commercial frontier labs all participate. International equivalents (UK, EU) are converging on similar designs.

**Q: What if my vendor uses an open-weights model that did not go through the Institute?**
A: You inherit more of the evaluation burden. Ask the vendor whether they have run equivalent in-house evals against the framework, and whether they will share results.

---

Source: https://callsphere.ai/blog/tw26w19-ai-safety-institute-pre-launch-evaluation-framework
