Cheaper, Faster Claude Agents: Caching & Batching
Cut Claude agent cost and latency with prompt caching, batching, and model routing. Practical levers, code, and a five-step plan to keep runs cheap.
Agentic AI, LLM engineering, and the models behind modern automation — multi-agent systems, LLM evaluation and comparisons, RAG, fine-tuning, AI infrastructure, security, and production AI engineering.
From the blog
Cut Claude agent cost and latency with prompt caching, batching, and model routing. Practical levers, code, and a five-step plan to keep runs cheap.
A field guide to debugging Claude agent failures — loops, wrong tool calls, and hallucinated arguments — with traces, guards, and a fix workflow.
What to put in Claude's context and what to leave out when classifying AI work — taxonomy injection, summaries over transcripts, and AEI-style accuracy trade-offs.
Connect MCP servers to an AEI-style Claude analytics agent the right way — auth kept server-side, narrow schemas, structured errors, and idempotent writes.
Reusable code-level patterns for classifying AI work with Claude — closed-vocabulary prompts, schema enforcement, context economy, and eval harnesses, AEI-style.
Step-by-step: classify agent conversations against a task taxonomy with Claude, label augment vs. automate, and aggregate safely — like the Anthropic Economic Index.
A technical teardown of how the Anthropic Economic Index turns private Claude usage into labor-market signal — classifier, taxonomy, privacy wall, and roll-ups.
Standardized Test Cases to Assess AI Model Performance
How Do You Really Know If Your LLM Is Good Enough? A Guide to Controlled Evaluation Metrics