Introducing Breeze TTS 2

Design any voice. Direct every performance. In real time.

Introducing Breeze TTS 2 — our new flagship voice model

THE SHIFT

From Content Creation to Real-Time Voice AI Experiences

The first wave of generative AI transformed content creation. Text, images, video, and speech can now be produced on demand, turning ideas into content with a prompt. Most of this content still takes the form of a finished artifact, created once and then consumed. The next wave brings interactivity, with real-time voice AI powering characters, games, companions, interactive stories, and voice agents that respond continuously to each user, context, and moment.

Content now responds continuously to the people experiencing it.

This paradigm shift expands what users expect from text to speech AI voices beyond naturalness. Natural-sounding speech remains the foundation of a high-quality voice experience, while interactive experiences introduce three additional requirements:

Diversity

Give every character a distinct voice.

Controllability

Adapt each performance as intent and context evolve.

Low Latency

Preserve the natural rhythm of interaction.

Together, these capabilities allow a voice to be designed for the character, directed for the moment, and delivered in real time.


THE NEW MODEL

Breeze TTS 2, a new speech model built for interactive experiences.

Breeze TTS 2 brings natural-language voice design, voice direction for any voice, and low-latency streaming generation together in a single model. Built for the next generation of real-time voice AI, it enables users to design any voice, direct every performance, and bring it to life in real time.

Advancing these capabilities also requires better ways to measure them. TTS evaluation has historically centered on naturalness, leaving voice design, voice direction, and latency less systematically explored. Breeze TTS 2 brings these capabilities together in a single model built for interactive voice experiences, giving creators and developers greater control over how voices are designed, performed, and delivered in real time. To support continued progress in these areas, we are also releasing open-source benchmarks for Voice Design, Voice Direction, and Latency. Now, let's take a closer look at the four capabilities that define Breeze TTS 2: voice design, voice direction, low latency, and multilingual speech.


CAPABILITY 01

Voice Design - the right AI voice for every character

Breeze TTS 2 gives every character a voice of their own, from a single protagonist to thousands of NPCs. Instead of relying only on preset voice libraries, creators can describe the character they have in mind and generate distinctive text to speech AI voices that better match each role, personality, and world.

#1 on the TTS Voice Design Benchmark
RankModel
Role FitVoice Diversity
1Breeze TTS 2
78.02708
2MiMo-V2.5-TTS
72.78202
3Inworld Voice Design
64.61378
4Fun-AudioGen-VD
63.82136
5Eleven v3
58.88111
6Qwen3-TTS-VD
58.4883
7VoxCPM2
55.28509

Breeze TTS 2 ranks first on the TTS Voice Design Benchmark, achieving the highest Role Fit score (78.02) and greatest Voice Diversity (708), with a 98.5% Transcript Pass rate. It leads the next-best model by 5.24 points in Role Fit and produces 39% more distinct voices than the closest competitor.

Animation & Comics

PROMPT

Masculine, mature adult. Massive heavy-set alien. Deep, gravelly, resonant bass-baritone. Gruff, authoritative, commanding, drill-sergeant cadence. Fast, high-intensity. Rough, weathered, immense power, stern mentorship.

SCRIPT

Chest up, eyes forward. Space does not care if you're tired, scared, or homesick; it only rewards discipline. Stay on my count, recruit, and I'll drag a champion out of you yet.

0:11

Film & TV

PROMPT

Young adult female, mid-20s, American accent. Professional broadcast journalist turning frantic. Clear, articulate but strained by panic. High tension, rapid pace, breathless. Polished to raw desperation.

SCRIPT

This is Elena Vance, reporting live from the east block where the crowd just broke the barricade—if anyone can hear me, we need medical teams now, please, they're getting closer.

0:08

Games

PROMPT

Female, young adult, ethereal trickster. High-pitched, breathy, mischievous. Fast, rhythmic delivery with giggles. Mocking sing-song riddles. Playful, taunting, otherworldly.

SCRIPT

Oh, you heard the whisper too? Clever little heartbeat. Follow the violet shimmer if you dare, but mind your shadow, love. It tends to trade secrets when the veil grows thin.

0:11

Literature & Stage

PROMPT

Male, mid-20s to early 30s. Smooth velvety tenor-range speaking voice with a light rasp. Soulful and deeply emotive, shifting from weary struggle to optimistic conviction. Impeccable diction, gentle Southern lilt, deliberate pace, and broad expressive pitch arcs.

SCRIPT

I've walked through too many long nights to mistake silence for peace, so when I speak of tomorrow, believe me, brother, I mean a dawn we can finally stand in together.

0:11

Myth & Folklore

PROMPT

Mature nomadic storyteller. Weathered, resonant speaking voice with rhythmic pacing and expressive pitch contours. Wide dynamic range from hushed whispers to booming proclamations; soulful.

SCRIPT

Come closer to the fire, child; the wind is telling old truths tonight. Hear how it bends the grass, then remember this: every road you take is also taking measure of you.

0:12

Research & News

PROMPT

Mature male, 50s-60s, Japanese-accented English. Dignified, measured, stoic tone. Clear texture with slight age rasp. Deliberate pacing, significant pauses. Conveys wisdom, authority, and visionary caution.

SCRIPT

We do not rebuild a nation with noise or haste. We place each stone with care, and when doubt rises, we answer it with discipline, patience, and an unshaken sense of duty.

0:15

Web & Periodical

PROMPT

Ancient male, massive scale. Deep, resonant bass, gravelly, subterranean. Slow, deliberate, aristocratic. Terrifying power, sibilant, predatory intelligence.

SCRIPT

Careful where you place your little hands, treasure-seeker; every coin in this vault knows its master, and I am not inclined to forgive even the smallest theft.

0:16
Explore 15,000+ character voices in the Voice Galaxy →

CAPABILITY 02

Voice Direction - the right performance for every moment

Use natural-language instructions and inline tags to shape emotion, intent, pace, and delivery while keeping every character recognizable. This gives creators and developers greater control over how an AI voice performs from one moment to the next without having to redesign the voice itself.

#1 on the TTS Voice Direction Benchmark
RankModel
Voice Direction ScoreSPK_SIM
1Breeze TTS 2
4.250.67
2MiMo-v2.5-TTS
3.760.66
3StepAudio 2.5 TTS
3.480.64
4Qwen-Audio 3.0-TTS-Plus
3.330.65
5Inworld TTS-2
3.010.71
6VoxCPM2
2.490.70

Breeze TTS 2 achieves the highest Voice Direction Score (4.25), outperforming the next-best model by 13%. Combined with a SIM score of 0.67, the result demonstrates its ability to follow natural-language direction across emotion, intent, pace, and delivery while keeping the selected voice recognizable across different performances.

Communicative Intent

PROMPT

Speak with an alluring, hushed intensity, using a hypnotic and persuasive rhythm to draw the listener in and convince them to acquire a rare item.

SCRIPT

You feel the energy radiating from it, don't you? This isn't just a piece of quartz. It's a conduit. For a small offering, you can take it home and finally clear that dark cloud hanging over your future. Trust me, you need this.

REFERENCE

0:20

OUTPUT

0:13

Composition

PROMPT

Speaking while heavily out of breath and gasping for air, trying to reassure the listener that everything is fine.

SCRIPT

[gasps] I'm okay... [pants] just give me a second to catch my breath. [gasps] I promise, I'm completely fine... just ran a little too fast.

REFERENCE

0:19

OUTPUT

0:15

Emotion

PROMPT

The speaker sounds acerbic and bitter, delivering the lines with heavy, passive-aggressive sarcasm.

SCRIPT

Enjoy your vacation while I stay here and clean up the mess you left behind for me. Have a wonderful time.

REFERENCE

0:31

OUTPUT

0:06

Role

PROMPT

Speak with a gritty, swaggering, and theatrical tone, projecting an aggressive and boastful attitude of a seafaring marauder.

SCRIPT

Drop the anchor and bring out the gold! If any of you scallywags think about crossing me, you'll be swimming with the sharks before sunset. This ocean belongs to me!

REFERENCE

0:28

OUTPUT

0:11

Variation

PROMPT

Start with a confident, reassuring tone. At the word 'Wait', suddenly shift to terrified, breathless alarm as if reacting to a frightening noise.

SCRIPT

There's absolutely nothing to worry about, I've done this a hundred times before. Wait... did you hear that sound? Something's definitely not right here.

REFERENCE

0:16

OUTPUT

0:09
Direct a voice of your own for free in BreezeBlue Creator →

CAPABILITY 03

Low Latency - Real-Time Voice AI at the Speed of Conversation

Breeze TTS 2 brings naturalness, voice design, and voice direction into a single low-latency model built for real-time voice AI applications. Fast time to first audio helps conversations stay responsive and preserves the natural rhythm of interactive voice experiences.

#1 on the TTS Latency Benchmark
0ms250ms500ms750ms1000msTTFB p50TTFA p50TTFA p95Breeze TTS 2Breeze TTS 2 — TTFB p50 119.4ms, TTFA p50 133.6ms, TTFA p95 163.3msElevenLabs Flash v2.5ElevenLabs Flash v2.5 — TTFB p50 134.8ms, TTFA p50 154.5ms, TTFA p95 207.6msFish Audio S2.1 ProFish Audio S2.1 Pro — TTFB p50 173.6ms, TTFA p50 175.3ms, TTFA p95 271.4msInworld TTS-2Inworld TTS-2 — TTFB p50 163.1ms, TTFA p50 189.4ms, TTFA p95 220.4msCartesia Sonic 3.5Cartesia Sonic 3.5 — TTFB p50 106.8ms, TTFA p50 241.9ms, TTFA p95 344.2msxAI TTSxAI TTS — TTFB p50 249.1ms, TTFA p50 318.3ms, TTFA p95 361.1msSpeechify Simba 3.2Speechify Simba 3.2 — TTFB p50 374.5ms, TTFA p50 379.0ms, TTFA p95 424.5msAsync Flash 1.5Async Flash 1.5 — TTFB p50 247.0ms, TTFA p50 438.1ms, TTFA p95 479.6msElevenLabs v3ElevenLabs v3 — TTFB p50 544.0ms, TTFA p50 628.7ms, TTFA p95 917.9ms

Breeze TTS 2 ranks first on the TTS Latency Benchmark, delivering the fastest time to first audio at both p50 and p95. Its consistently low latency keeps realtime interactions responsive and preserves the natural rhythm of conversation.

Streaming Breeze TTS 2 in realtime

The realtime API keeps one WebSocket session open across conversation turns: append text as it arrives, and raw PCM audio streams back while the model is still generating.

import { BreezeBlueClient } from "@breeze.blue/sdk";

const client = new BreezeBlueClient();

// One WebSocket session, many conversation turns.
const connection = await client.textToSpeech.realtime.connect("voc_...", {
  modelId: "breeze-tts-2",
});

const consumer = (async () => {
  for await (const message of connection) {
    if (message.type === "audio") {
      play(message.audio); // raw PCM, streamed while the model generates
    } else if (message.type === "turn.done") {
      return;
    }
  }
})();

connection.startTurn("turn_1");
connection.appendText("Hello from Breeze TTS 2.");
connection.flush();
connection.endTurn();

await consumer;
connection.close();
Get your BreezeBlue API key →

CAPABILITY 04

Multilingual Speech

Create natural, expressive text to speech AI voices across 50 languages, with accent control for adapting delivery to different regions and audiences.

Choose a language

Every journey finds its meaning when someone dares to take the first step.

Original: Every journey finds its meaning when someone dares to take the first step.

Bring your script to life in BreezeBlue Creator →

BreezeBlue Creator

Design, direct, and generate voices in one browser-based creator tool.

BreezeBlue Creator text-to-speech workspace

Text to Speech API and SDKs

Integrate BreezeBlue Text to Speech into your product via APIs or SDKs.

BreezeBlue Text to Speech API code sample

Enterprise Solution

Bring production-grade voice experiences to your organization

Enterprise voice experience illustration