Deploy a Voice Agent on Modal with Python and Serverless GPU
Modal turns a Python function into autoscaling serverless compute with optional GPU. Deploy a LiveKit Agent with one command and get pay-per-second billing.
Agentic AI, LLM engineering, and the models behind modern automation — multi-agent systems, LLM evaluation and comparisons, RAG, fine-tuning, AI infrastructure, security, and production AI engineering.
From the blog
Modal turns a Python function into autoscaling serverless compute with optional GPU. Deploy a LiveKit Agent with one command and get pay-per-second billing.
Confidential computing, secure enclaves, and zero-knowledge proofs are no longer research curiosities. Here is how a 2026 HIPAA-aligned AI voice platform uses private compute on PHI.
Run Whisper, Kokoro, and LFM2.5-Audio entirely in the browser with ONNX Runtime Web + WebGPU. Flash Attention, qMoE, sub-100ms latency on a laptop. Privacy-first voice without a backend.
Memory for Inference: Why Serving LLMs Is Really a Memory Problem
Where agentic AI is heading next — longer autonomy, multi-agent norms, richer MCP and skills — and how to prepare, from a Built-with-Opus hackathon.
The metrics and signals that prove agentic AI works — success rate, cost, latency, and intervention rate — from a Built-with-Opus Claude Code hackathon.
A realistic end-to-end agentic AI walkthrough — from vague problem to shipped, verified feature with Claude Code, MCP, and Skills at a hackathon.
Failure scenarios, blast radius, and containment patterns for agentic AI — lessons from a Built-with-Opus Claude Code hackathon on running agents safely.
A Built-with-Opus Claude Code hackathon revealed which skills, roles, and hiring shifts make agentic AI work. Here's what engineering teams should learn.