From Trace to Production Fix: An End-to-End Observability Workflow for Agents
A real workflow: user complaint → LangSmith trace → reproduce in dataset → fix → ship → re-eval. Principal-engineer notes, real numbers, honest tradeoffs.
Agentic AI, LLM engineering, and the models behind modern automation — multi-agent systems, LLM evaluation and comparisons, RAG, fine-tuning, AI infrastructure, security, and production AI engineering.
From the blog
A real workflow: user complaint → LangSmith trace → reproduce in dataset → fix → ship → re-eval. Principal-engineer notes, real numbers, honest tradeoffs.
Anthropic unveiled 10 pre-built finance agent templates on May 5, 2026 across pitchbook building, KYC screening, and month-end close. What each template does and the hours it replaces.
Anthropic shipped finance plugins for Claude Cowork and Claude Code on May 5, 2026. How analysts use them in practice and what the plugin model means for adoption.
Anthropic's restricted Mythos model is reshaping vuln discovery. Inside the Mozilla Firefox case, what it means for AppSec, and where voice AI fits.
The full metric set for evaluating production voice agents — STT word error rate, end-to-end latency budgets, RAG grounding, prosody, and the metrics that actually correlate with retention.
Where agentic AI in banking is heading next, from longer-running agents to agent-to-agent MCP workflows, and how to prepare your team now.
Eval pass rate, override rate, citation validity, cycle time, and cost per case: the metrics that prove a Claude agent works in banking and fintech.
A realistic step-by-step build of a Claude exception-triage agent for lenders, from messy problem to shipped, audited, monitored outcome.
Failure modes, blast radius, and containment patterns for deploying Claude agents safely across banking, lending, and fintech workflows.