Apresentamos o Breeze TTS 2

Desenhe qualquer voz. Dirija cada performance. Em tempo real.

Apresentamos o Breeze TTS 2, nosso novo modelo de voz principal

A MUDANÇA

Da criação de conteúdo às experiências de voz com IA em tempo real

A primeira onda da IA generativa transformou a criação de conteúdo. Texto, imagens, vídeo e fala já podem ser produzidos sob demanda, e um prompt basta para transformar uma ideia em conteúdo. Ainda assim, a maior parte desse conteúdo continua sendo uma peça acabada: criada uma vez e depois consumida. A próxima onda traz interatividade — com IA de voz em tempo real dando vida a personagens, jogos, companheiros virtuais, histórias interativas e agentes de voz que respondem continuamente a cada usuário, cada contexto e cada momento.

Agora o conteúdo responde continuamente a quem o está vivenciando.

Essa mudança de paradigma amplia o que as pessoas esperam das vozes de IA de texto para voz para além da naturalidade. Uma fala com som natural continua sendo a base de uma experiência de voz de qualidade, mas as experiências interativas trazem três exigências adicionais:

Diversidade

Dar a cada personagem uma voz inconfundível.

Controlabilidade

Adaptar cada performance conforme a intenção e o contexto mudam.

Baixa latência

Preservar o ritmo natural da interação.

Juntas, essas capacidades permitem que uma voz seja desenhada para o personagem, dirigida para o momento e entregue em tempo real.


O NOVO MODELO

Breeze TTS 2: um novo modelo de fala feito para experiências interativas.

O Breeze TTS 2 reúne em um único modelo o design de voz por linguagem natural, a direção de voz para qualquer voz e a geração em streaming de baixa latência. Feito para a próxima geração de IA de voz em tempo real, ele permite desenhar qualquer voz, dirigir cada performance e trazê-la à vida em tempo real.

Avançar nessas capacidades também exige formas melhores de medi-las. A avaliação de TTS sempre girou em torno da naturalidade, deixando design de voz, direção de voz e latência pouco explorados de forma sistemática. O Breeze TTS 2 reúne essas capacidades em um único modelo criado para experiências de voz interativas, dando a criadores e desenvolvedores mais controle sobre como as vozes são desenhadas, interpretadas e entregues em tempo real. Para sustentar o progresso nessas áreas, também estamos publicando benchmarks de código aberto de design de voz, direção de voz e latência. Agora, vamos olhar de perto as quatro capacidades que definem o Breeze TTS 2: design de voz, direção de voz, baixa latência e fala multilíngue.


CAPACIDADE 01

Design de voz: a voz de IA certa para cada personagem

De um único protagonista a milhares de NPCs, o Breeze TTS 2 dá a cada personagem uma voz própria. Em vez de depender apenas de bibliotecas de vozes prontas, os criadores descrevem o personagem que têm em mente e geram vozes de IA de texto para voz marcantes, mais alinhadas a cada papel, personalidade e universo.

Nº 1 no Benchmark de Design de voz em TTS
PosiçãoModelo
Role FitVoice Diversity
1Breeze TTS 2
78.02708
2MiMo-V2.5-TTS
72.78202
3Inworld Voice Design
64.61378
4Fun-AudioGen-VD
63.82136
5Eleven v3
58.88111
6Qwen3-TTS-VD
58.4883
7VoxCPM2
55.28509

O Breeze TTS 2 lidera o Benchmark de Design de voz em TTS: tem a maior pontuação de Role Fit (78.02) e a maior Voice Diversity (708), com 98.5% de aprovação em Transcript Pass. Em Role Fit, supera o segundo colocado por 5.24 pontos, e gera 39% mais vozes distintas que o concorrente mais próximo.

Animation & Comics

PROMPT

Masculine, mature adult. Massive heavy-set alien. Deep, gravelly, resonant bass-baritone. Gruff, authoritative, commanding, drill-sergeant cadence. Fast, high-intensity. Rough, weathered, immense power, stern mentorship.

SCRIPT

Chest up, eyes forward. Space does not care if you're tired, scared, or homesick; it only rewards discipline. Stay on my count, recruit, and I'll drag a champion out of you yet.

0:11

Film & TV

PROMPT

Young adult female, mid-20s, American accent. Professional broadcast journalist turning frantic. Clear, articulate but strained by panic. High tension, rapid pace, breathless. Polished to raw desperation.

SCRIPT

This is Elena Vance, reporting live from the east block where the crowd just broke the barricade—if anyone can hear me, we need medical teams now, please, they're getting closer.

0:08

Games

PROMPT

Female, young adult, ethereal trickster. High-pitched, breathy, mischievous. Fast, rhythmic delivery with giggles. Mocking sing-song riddles. Playful, taunting, otherworldly.

SCRIPT

Oh, you heard the whisper too? Clever little heartbeat. Follow the violet shimmer if you dare, but mind your shadow, love. It tends to trade secrets when the veil grows thin.

0:11

Literature & Stage

PROMPT

Male, mid-20s to early 30s. Smooth velvety tenor-range speaking voice with a light rasp. Soulful and deeply emotive, shifting from weary struggle to optimistic conviction. Impeccable diction, gentle Southern lilt, deliberate pace, and broad expressive pitch arcs.

SCRIPT

I've walked through too many long nights to mistake silence for peace, so when I speak of tomorrow, believe me, brother, I mean a dawn we can finally stand in together.

0:11

Myth & Folklore

PROMPT

Mature nomadic storyteller. Weathered, resonant speaking voice with rhythmic pacing and expressive pitch contours. Wide dynamic range from hushed whispers to booming proclamations; soulful.

SCRIPT

Come closer to the fire, child; the wind is telling old truths tonight. Hear how it bends the grass, then remember this: every road you take is also taking measure of you.

0:12

Research & News

PROMPT

Mature male, 50s-60s, Japanese-accented English. Dignified, measured, stoic tone. Clear texture with slight age rasp. Deliberate pacing, significant pauses. Conveys wisdom, authority, and visionary caution.

SCRIPT

We do not rebuild a nation with noise or haste. We place each stone with care, and when doubt rises, we answer it with discipline, patience, and an unshaken sense of duty.

0:15

Web & Periodical

PROMPT

Ancient male, massive scale. Deep, resonant bass, gravelly, subterranean. Slow, deliberate, aristocratic. Terrifying power, sibilant, predatory intelligence.

SCRIPT

Careful where you place your little hands, treasure-seeker; every coin in this vault knows its master, and I am not inclined to forgive even the smallest theft.

0:16
Explore 15.000+ vozes de personagens na Voice Galaxy →

CAPACIDADE 02

Direção de voz: a performance certa para cada momento

Use instruções em linguagem natural e tags inline para moldar emoção, intenção, ritmo e entrega mantendo cada personagem reconhecível. Assim, criadores e desenvolvedores ganham mais controle sobre como uma voz de IA atua de um momento para o outro, sem precisar redesenhar a voz.

Nº 1 no Benchmark de Direção de voz em TTS
PosiçãoModelo
Pontuação de Direção de vozSPK_SIM
1Breeze TTS 2
4.250.67
2MiMo-v2.5-TTS
3.760.66
3StepAudio 2.5 TTS
3.480.64
4Qwen-Audio 3.0-TTS-Plus
3.330.65
5Inworld TTS-2
3.010.71
6VoxCPM2
2.490.70

O Breeze TTS 2 alcança a maior pontuação de direção de voz (4.25), 13% acima do segundo melhor modelo. Somado a um SIM de 0.67, o resultado mostra sua capacidade de seguir direção em linguagem natural quanto a emoção, intenção, ritmo e entrega, mantendo a voz escolhida reconhecível em performances diferentes.

Communicative Intent

PROMPT

Speak with an alluring, hushed intensity, using a hypnotic and persuasive rhythm to draw the listener in and convince them to acquire a rare item.

SCRIPT

You feel the energy radiating from it, don't you? This isn't just a piece of quartz. It's a conduit. For a small offering, you can take it home and finally clear that dark cloud hanging over your future. Trust me, you need this.

REFERENCE

0:20

OUTPUT

0:13

Composition

PROMPT

Speaking while heavily out of breath and gasping for air, trying to reassure the listener that everything is fine.

SCRIPT

[gasps] I'm okay... [pants] just give me a second to catch my breath. [gasps] I promise, I'm completely fine... just ran a little too fast.

REFERENCE

0:19

OUTPUT

0:15

Emotion

PROMPT

The speaker sounds acerbic and bitter, delivering the lines with heavy, passive-aggressive sarcasm.

SCRIPT

Enjoy your vacation while I stay here and clean up the mess you left behind for me. Have a wonderful time.

REFERENCE

0:31

OUTPUT

0:06

Role

PROMPT

Speak with a gritty, swaggering, and theatrical tone, projecting an aggressive and boastful attitude of a seafaring marauder.

SCRIPT

Drop the anchor and bring out the gold! If any of you scallywags think about crossing me, you'll be swimming with the sharks before sunset. This ocean belongs to me!

REFERENCE

0:28

OUTPUT

0:11

Variation

PROMPT

Start with a confident, reassuring tone. At the word 'Wait', suddenly shift to terrified, breathless alarm as if reacting to a frightening noise.

SCRIPT

There's absolutely nothing to worry about, I've done this a hundred times before. Wait... did you hear that sound? Something's definitely not right here.

REFERENCE

0:16

OUTPUT

0:09
Dirija sua própria voz de graça no BreezeBlue Creator →

CAPACIDADE 03

Baixa latência: IA de voz em tempo real na velocidade da conversa

O Breeze TTS 2 reúne naturalidade, design de voz e direção de voz em um único modelo de baixa latência, feito para aplicações de IA de voz em tempo real. Um tempo curto até o primeiro áudio mantém a conversa responsiva e preserva o ritmo natural das experiências de voz interativas.

Nº 1 no Benchmark de latência em TTS
0ms250ms500ms750ms1000msTTFB p50TTFA p50TTFA p95Breeze TTS 2Breeze TTS 2 — TTFB p50 119.4ms, TTFA p50 133.6ms, TTFA p95 163.3msElevenLabs Flash v2.5ElevenLabs Flash v2.5 — TTFB p50 134.8ms, TTFA p50 154.5ms, TTFA p95 207.6msFish Audio S2.1 ProFish Audio S2.1 Pro — TTFB p50 173.6ms, TTFA p50 175.3ms, TTFA p95 271.4msInworld TTS-2Inworld TTS-2 — TTFB p50 163.1ms, TTFA p50 189.4ms, TTFA p95 220.4msCartesia Sonic 3.5Cartesia Sonic 3.5 — TTFB p50 106.8ms, TTFA p50 241.9ms, TTFA p95 344.2msxAI TTSxAI TTS — TTFB p50 249.1ms, TTFA p50 318.3ms, TTFA p95 361.1msSpeechify Simba 3.2Speechify Simba 3.2 — TTFB p50 374.5ms, TTFA p50 379.0ms, TTFA p95 424.5msAsync Flash 1.5Async Flash 1.5 — TTFB p50 247.0ms, TTFA p50 438.1ms, TTFA p95 479.6msElevenLabs v3ElevenLabs v3 — TTFB p50 544.0ms, TTFA p50 628.7ms, TTFA p95 917.9ms

O Breeze TTS 2 lidera o Benchmark de latência em TTS, com o menor tempo até o primeiro áudio tanto em p50 quanto em p95. Uma latência baixa e consistente mantém as interações em tempo real ágeis e preserva o ritmo natural da conversa.

Breeze TTS 2 em streaming, em tempo real

A API realtime mantém uma única sessão WebSocket aberta ao longo dos turnos da conversa: você acrescenta o texto conforme ele chega e o áudio PCM bruto volta em streaming enquanto o modelo ainda está gerando.

import { BreezeBlueClient } from "@breeze.blue/sdk";

const client = new BreezeBlueClient();

// One WebSocket session, many conversation turns.
const connection = await client.textToSpeech.realtime.connect("voc_...", {
  modelId: "breeze-tts-2",
});

const consumer = (async () => {
  for await (const message of connection) {
    if (message.type === "audio") {
      play(message.audio); // raw PCM, streamed while the model generates
    } else if (message.type === "turn.done") {
      return;
    }
  }
})();

connection.startTurn("turn_1");
connection.appendText("Hello from Breeze TTS 2.");
connection.flush();
connection.endTurn();

await consumer;
connection.close();
Obtenha sua API key da BreezeBlue →

CAPACIDADE 04

Fala multilíngue

Crie vozes de IA de texto para voz naturais e expressivas em 50 idiomas, com controle de sotaque para adaptar a entrega a diferentes regiões e públicos.

Escolha um idioma

Every journey finds its meaning when someone dares to take the first step.

Original: Every journey finds its meaning when someone dares to take the first step.

Dê vida ao seu roteiro no BreezeBlue Creator →

BreezeBlue Creator

Desenhe, dirija e gere vozes em uma única ferramenta de criação que roda no navegador.

BreezeBlue Creator text-to-speech workspace

API e SDKs de Texto para voz

Integre o Texto para voz da BreezeBlue ao seu produto por APIs ou SDKs.

BreezeBlue Text to Speech API code sample

Solução Enterprise

Leve experiências de voz de nível de produção para a sua organização

Enterprise voice experience illustration