Chunked Streaming TTS: Time-to-First-Audio Optimization (2026)
ElevenLabs Flash v2.5 hits 75ms inference; Cartesia Sonic streams under 100ms. We tune chunk_length_schedule, sentence-boundary streaming, and TTFB to keep TTS under 200ms of the voice budget.