


By Sagar Shankaran, Founder of CallSphere
Text to speech with emotion in 2026 means dynamic prosody, real anger, real warmth — not robotic voices. Here is how it works and what voice agents need.
Key takeaways
This is part of our Siri Voice Generator guide.
Text to speech with emotion in 2026 means a TTS system that can speak the same words with materially different prosody, pacing, pitch, and tonal warmth based on a directive — either a natural-language instruction or a structured emotion tag. The output sounds like a person who is actually feeling something, not a flat voice with pitch contour bolted on.
This is a real shift. As recently as 2023, "emotional TTS" was mostly SSML hacks — phoneme-level pitch adjustments and rate changes layered on a concatenative engine. The result sounded uncanny. In 2026, the leading models — OpenAI's GPT-Realtime-2 stack, ElevenLabs v3, Cartesia Sonic-2 — generate audio end-to-end and accept emotion directives in plain English. You write "speak this warmly and slowly, as if reassuring a worried patient" and the model does it.
I use this every day in CallSphere. The healthcare agent uses a warm, slow voice tuned for elderly callers. The sales outbound agent uses an upbeat, confident voice. The after-hours emergency escalation agent uses a calm, reassuring voice that drops the caller's heart rate, not raises it. None of those voices use the same prosody preset.
GPT-Realtime-2, which OpenAI shipped on May 7, 2026, embeds emotion steering directly into the system prompt and the per-turn instruction. The model accepts directives like "speak with quiet enthusiasm" or "convey gentle empathy" and adjusts prosody, pacing, and pitch contour automatically. There is no separate emotion tag layer.
The pricing matters for production teams: audio input at $32 per 1M tokens, audio output at $64 per 1M tokens, cached input at $0.40 per 1M tokens. If your system prompt includes a 600-token emotion specification, cached input makes it negligible across thousands of calls.
The 128K context window means you can include long emotion playbooks — "if the caller raises their voice, switch to a calm, low register; if the caller jokes, match the warmth" — without trimming. We do exactly this in CallSphere's prompts.
The speech to text side matters as much as the TTS side. The agent needs to hear the caller's emotion to respond appropriately. The leading 2026 models for this are GPT-Realtime-Whisper (OpenAI), Deepgram Nova-3, and AssemblyAI Universal-2. All three transcribe with paralinguistic features — pace, volume, hesitation markers — which the agent's logic layer uses to choose the next emotion directive.
CallSphere uses GPT-Realtime-Whisper for STT and GPT-Realtime-2 for TTS. The pipeline runs in roughly 600ms turn latency on a 4G connection, which is the floor for natural conversation. Slower than 800ms and callers start to interrupt.
For specific voice profiles — woman text to speech, australian text to speech, adam text to speech, text to speech with characters — the new generative models render any voice profile from a sample or a description. CallSphere ships with 60+ voice profiles across 57+ languages, and you can clone a custom voice from a 30-second sample if your contract allows it.
Hear it before you finish reading
Talk to a live CallSphere AI voice agent in your browser — 60 seconds, no signup.
You write a natural-language directive. Examples we use in production:
For accent-specific use cases — australian text to speech, japanese text to speech (text to speech japanese), british english — you specify the accent and locale in the voice config. The model renders it natively. You do not need a separate accent layer.
Our TTS stack runs on GPT-Realtime-2 with per-agent prompts:
Hear all 6 CallSphere voices in the live demo →
A dental group in Boston deployed CallSphere's healthcare agent in March 2026 for inbound appointment booking and after-hours triage. The first week, their highest-volume call type was elderly patients calling about appointment confirmations.
We tuned the agent's voice to a warm, slow, lower-register profile based on patient feedback. Average call duration on confirmation calls dropped from 3:40 to 2:10 because the agent's pacing matched the caller's. CSAT on those calls went up 18 points. The clinic's office manager called it "the first AI that sounds like it actually cares."
The mechanism is just prompt-level emotion steering: "Speak warmly, slowly, and with the patience of a 30-year veteran nurse." The model rendered that into prosody. No SSML, no custom voice training, no $50K voice budget.
Text to speech with emotion is included in every CallSphere plan:
7-day free pilot, no credit card. 3-5 day setup.
What is text to speech with emotion in 2026?
Text to speech with emotion in 2026 is generative audio that adjusts prosody, pacing, pitch, and warmth based on a directive — either a natural-language instruction ("speak warmly and slowly") or a structured emotion tag ("calm", "enthusiastic", "empathetic"). The output is rendered end-to-end by a generative model like GPT-Realtime-2, not assembled from phoneme samples. The result sounds like a person with a mood, not a flat voice with pitch contour applied.
Still reading? Stop comparing — try CallSphere live.
CallSphere ships complete AI voice agents per industry — 14 tools for healthcare, 10 agents for real estate, 4 specialists for salons. See how it actually handles a call before you book a demo.
What is the best text to speech model for a voice agent?
For voice agents in production, the best text to speech model in 2026 is GPT-Realtime-2 from OpenAI, with ElevenLabs v3 and Cartesia Sonic-2 as strong alternatives. GPT-Realtime-2 is the default at CallSphere because it integrates STT, reasoning, and TTS in a single model, which keeps turn latency under 800ms. ElevenLabs is the best for ultra-realistic single-voice work; Cartesia is the best for low-latency streaming use cases.
Can I use woman text to speech voices with emotion in CallSphere?
Yes. CallSphere ships 30+ female voices across 57+ languages with full emotion steering — warm, calm, confident, enthusiastic, professional. You pick the base voice in the admin console and add emotion directives in the agent prompt. Custom voice cloning is available on the Scale tier ($1,499/mo) with a 30-second sample and contractual voice licensing.
Is there a good adam text to speech voice for sales calls?
Yes. "Adam" is a common male voice profile name across ElevenLabs and other TTS engines. CallSphere ships multiple male voice options including a confident American male profile that we use for our default sales outbound agent. You can also clone a custom Adam-style voice on the Scale tier from a 30-second sample.
How does text to speech with characters work in 2026?
Text to speech with characters — meaning multiple distinct voice personas in one conversation or piece of content — is straightforward in 2026 generative models. You specify the voice profile per turn or per character tag. CallSphere supports this for hold messages, multi-language IVR replacement, and any use case where the agent needs to play multiple roles in one call. Most use cases are well served by a single voice with emotion steering across turns rather than multiple character voices.
What is the best speech to text model for paralinguistic features?
For paralinguistic features — pace, volume, hesitation, emotion in the caller's speech — the leading 2026 speech to text models are GPT-Realtime-Whisper, Deepgram Nova-3, and AssemblyAI Universal-2. CallSphere uses GPT-Realtime-Whisper because it integrates natively with GPT-Realtime-2 and shares the 128K context window, which lets the agent reason about caller emotion across the full conversation.
Can I get australian text to speech with emotion in CallSphere?
Yes. Australian English is one of the 57+ language and accent options. The same emotion steering directives work — warm, calm, confident, enthusiastic — and the model renders them in the Australian accent. We have customers using Australian voices for sales outbound to Australian markets and for hospitality concierge in Australian hotels.
How much does emotion-aware TTS cost in 2026?
GPT-Realtime-2's audio pricing is $32 per 1M input tokens, $64 per 1M output tokens, and $0.40 per 1M cached input tokens. A 10-minute voice call typically costs $0.30-$0.60 in raw model spend. CallSphere bundles this into our flat-rate plans ($149-$1,499/mo), so you do not see token-level pricing — you see interaction counts. A "Growth" customer at $499/mo gets 10,000 interactions including all emotion-aware TTS at no upcharge.

Written by
Sagar Shankaran· Founder, CallSphere
LinkedInSagar Shankaran is the founder of CallSphere, where he builds production AI voice and chat agents deployed across healthcare, hospitality, real estate, and home services. He writes about agentic AI, LLM engineering, and shipping voice agents that handle real calls in production.
See how AI voice agents work for your industry. Live demo available -- no signup required.
Voice activated business systems in 2026 go beyond wake words. Here is how AI voice agents, TTS, and STT actually work together — and what to deploy.
IVR software is being replaced by AI voice agents, but not entirely. Here is when IVR still makes sense, when it does not, and how CallSphere handles both.
Sesame voice has emerged as a top TTS option in 2026. Here is what the sesame voice model actually is, how it sounds, and where CallSphere uses it.
Real conversational AI examples from production deployments in 2026. Healthcare, real estate, sales, salon, after-hours, and hotel use cases, with numbers.
Chatable AI is the new shorthand for conversational agents you can actually talk to. Here is what makes one chatable and how CallSphere ships them in 3-5 days.
An automated answering service in 2026 is an AI voice agent that books, qualifies, and escalates 24/7. Here is how it works and what it costs.
© 2026 CallSphere Inc. All rights reserved.
Made within San Francisco
Watch how CallSphere handles real customer calls, schedules appointments, and processes payments — live.