Presentamos Breeze TTS 2

Diseña cualquier voz. Dirige cada interpretación. En tiempo real.

Presentamos Breeze TTS 2, nuestro nuevo modelo de voz insignia

EL CAMBIO

De la creación de contenido a las experiencias de voz con IA en tiempo real

La primera ola de IA generativa transformó la creación de contenido. Hoy se pueden producir texto, imágenes, video y voz bajo demanda: basta un prompt para convertir una idea en contenido. Aun así, casi todo ese contenido sigue siendo una pieza terminada, creada una sola vez y luego consumida. La siguiente ola trae la interactividad: la voz con IA en tiempo real impulsa personajes, juegos, acompañantes, historias interactivas y agentes de voz que responden de forma continua a cada usuario, cada contexto y cada momento.

El contenido ahora responde de forma continua a quienes lo están viviendo.

Este cambio de paradigma amplía lo que los usuarios esperan de las voces de IA de texto a voz más allá de la naturalidad. Un habla que suene natural sigue siendo la base de una experiencia de voz de calidad, y las experiencias interactivas suman tres requisitos más:

Diversidad

Dale a cada personaje una voz inconfundible.

Controlabilidad

Adapta cada interpretación a medida que evolucionan la intención y el contexto.

Baja latencia

Conserva el ritmo natural de la interacción.

Juntas, estas capacidades permiten diseñar una voz para el personaje, dirigirla para el momento y entregarla en tiempo real.


EL NUEVO MODELO

Breeze TTS 2, un nuevo modelo de voz creado para experiencias interactivas.

Breeze TTS 2 reúne en un solo modelo el diseño de voz por lenguaje natural, la dirección de voz para cualquier voz y la generación por streaming de baja latencia. Pensado para la próxima generación de voz con IA en tiempo real, permite a los usuarios diseñar cualquier voz, dirigir cada interpretación y darle vida en tiempo real.

Llevar más lejos estas capacidades exige también mejores formas de medirlas. La evaluación de TTS se ha centrado históricamente en la naturalidad, y ha dejado el diseño de voz, la dirección de voz y la latencia mucho menos explorados de forma sistemática. Breeze TTS 2 reúne estas capacidades en un único modelo creado para experiencias de voz interactivas, y le da a creadores y desarrolladores mayor control sobre cómo se diseñan, se interpretan y se entregan las voces en tiempo real. Para impulsar el avance en estas áreas, publicamos además benchmarks de código abierto de Diseño de voz, Dirección de voz y Latencia. Veamos ahora más de cerca las cuatro capacidades que definen a Breeze TTS 2: diseño de voz, dirección de voz, baja latencia y habla multilingüe.


CAPACIDAD 01

Diseño de voz: la voz de IA adecuada para cada personaje

Breeze TTS 2 le da a cada personaje una voz propia, desde un único protagonista hasta miles de NPC. En lugar de depender solo de bibliotecas de voces predefinidas, los creadores pueden describir el personaje que tienen en mente y generar voces de IA de texto a voz distintivas, más afines a cada papel, cada personalidad y cada universo.

#1 en el Benchmark de diseño de voz TTS
PuestoModelo
Role FitVoice Diversity
1Breeze TTS 2
78.02708
2MiMo-V2.5-TTS
72.78202
3Inworld Voice Design
64.61378
4Fun-AudioGen-VD
63.82136
5Eleven v3
58.88111
6Qwen3-TTS-VD
58.4883
7VoxCPM2
55.28509

Breeze TTS 2 ocupa el primer puesto del Benchmark de diseño de voz TTS: logra la puntuación de Role Fit más alta (78.02) y la mayor Voice Diversity (708), con una tasa de Transcript Pass del 98.5%. Le saca 5.24 puntos de ventaja en Role Fit al segundo mejor modelo y produce un 39% más de voces distintas que su competidor más cercano.

Animation & Comics

PROMPT

Masculine, mature adult. Massive heavy-set alien. Deep, gravelly, resonant bass-baritone. Gruff, authoritative, commanding, drill-sergeant cadence. Fast, high-intensity. Rough, weathered, immense power, stern mentorship.

SCRIPT

Chest up, eyes forward. Space does not care if you're tired, scared, or homesick; it only rewards discipline. Stay on my count, recruit, and I'll drag a champion out of you yet.

0:11

Film & TV

PROMPT

Young adult female, mid-20s, American accent. Professional broadcast journalist turning frantic. Clear, articulate but strained by panic. High tension, rapid pace, breathless. Polished to raw desperation.

SCRIPT

This is Elena Vance, reporting live from the east block where the crowd just broke the barricade—if anyone can hear me, we need medical teams now, please, they're getting closer.

0:08

Games

PROMPT

Female, young adult, ethereal trickster. High-pitched, breathy, mischievous. Fast, rhythmic delivery with giggles. Mocking sing-song riddles. Playful, taunting, otherworldly.

SCRIPT

Oh, you heard the whisper too? Clever little heartbeat. Follow the violet shimmer if you dare, but mind your shadow, love. It tends to trade secrets when the veil grows thin.

0:11

Literature & Stage

PROMPT

Male, mid-20s to early 30s. Smooth velvety tenor-range speaking voice with a light rasp. Soulful and deeply emotive, shifting from weary struggle to optimistic conviction. Impeccable diction, gentle Southern lilt, deliberate pace, and broad expressive pitch arcs.

SCRIPT

I've walked through too many long nights to mistake silence for peace, so when I speak of tomorrow, believe me, brother, I mean a dawn we can finally stand in together.

0:11

Myth & Folklore

PROMPT

Mature nomadic storyteller. Weathered, resonant speaking voice with rhythmic pacing and expressive pitch contours. Wide dynamic range from hushed whispers to booming proclamations; soulful.

SCRIPT

Come closer to the fire, child; the wind is telling old truths tonight. Hear how it bends the grass, then remember this: every road you take is also taking measure of you.

0:12

Research & News

PROMPT

Mature male, 50s-60s, Japanese-accented English. Dignified, measured, stoic tone. Clear texture with slight age rasp. Deliberate pacing, significant pauses. Conveys wisdom, authority, and visionary caution.

SCRIPT

We do not rebuild a nation with noise or haste. We place each stone with care, and when doubt rises, we answer it with discipline, patience, and an unshaken sense of duty.

0:15

Web & Periodical

PROMPT

Ancient male, massive scale. Deep, resonant bass, gravelly, subterranean. Slow, deliberate, aristocratic. Terrifying power, sibilant, predatory intelligence.

SCRIPT

Careful where you place your little hands, treasure-seeker; every coin in this vault knows its master, and I am not inclined to forgive even the smallest theft.

0:16
Explora 15,000+ voces de personajes en Voice Galaxy →

CAPACIDAD 02

Dirección de voz: la interpretación adecuada para cada momento

Usa instrucciones en lenguaje natural y etiquetas en línea para moldear la emoción, la intención, el ritmo y la entrega, sin que ningún personaje deje de ser reconocible. Así, creadores y desarrolladores ganan control sobre cómo actúa una voz de IA de un momento a otro, sin tener que rediseñar la voz.

#1 en el Benchmark de dirección de voz TTS
PuestoModelo
Puntuación de Dirección de vozSPK_SIM
1Breeze TTS 2
4.250.67
2MiMo-v2.5-TTS
3.760.66
3StepAudio 2.5 TTS
3.480.64
4Qwen-Audio 3.0-TTS-Plus
3.330.65
5Inworld TTS-2
3.010.71
6VoxCPM2
2.490.70

Breeze TTS 2 alcanza la Puntuación de Dirección de voz más alta (4.25) y supera al segundo mejor modelo por un 13%. Sumado a una puntuación SIM de 0.67, el resultado demuestra su capacidad de seguir indicaciones en lenguaje natural sobre emoción, intención, ritmo y entrega, manteniendo la voz elegida reconocible en interpretaciones distintas.

Communicative Intent

PROMPT

Speak with an alluring, hushed intensity, using a hypnotic and persuasive rhythm to draw the listener in and convince them to acquire a rare item.

SCRIPT

You feel the energy radiating from it, don't you? This isn't just a piece of quartz. It's a conduit. For a small offering, you can take it home and finally clear that dark cloud hanging over your future. Trust me, you need this.

REFERENCE

0:20

OUTPUT

0:13

Composition

PROMPT

Speaking while heavily out of breath and gasping for air, trying to reassure the listener that everything is fine.

SCRIPT

[gasps] I'm okay... [pants] just give me a second to catch my breath. [gasps] I promise, I'm completely fine... just ran a little too fast.

REFERENCE

0:19

OUTPUT

0:15

Emotion

PROMPT

The speaker sounds acerbic and bitter, delivering the lines with heavy, passive-aggressive sarcasm.

SCRIPT

Enjoy your vacation while I stay here and clean up the mess you left behind for me. Have a wonderful time.

REFERENCE

0:31

OUTPUT

0:06

Role

PROMPT

Speak with a gritty, swaggering, and theatrical tone, projecting an aggressive and boastful attitude of a seafaring marauder.

SCRIPT

Drop the anchor and bring out the gold! If any of you scallywags think about crossing me, you'll be swimming with the sharks before sunset. This ocean belongs to me!

REFERENCE

0:28

OUTPUT

0:11

Variation

PROMPT

Start with a confident, reassuring tone. At the word 'Wait', suddenly shift to terrified, breathless alarm as if reacting to a frightening noise.

SCRIPT

There's absolutely nothing to worry about, I've done this a hundred times before. Wait... did you hear that sound? Something's definitely not right here.

REFERENCE

0:16

OUTPUT

0:09
Dirige tu propia voz gratis en BreezeBlue Creator →

CAPACIDAD 03

Baja latencia: voz con IA en tiempo real a la velocidad de la conversación

Breeze TTS 2 reúne naturalidad, diseño de voz y dirección de voz en un único modelo de baja latencia, creado para aplicaciones de voz con IA en tiempo real. Un tiempo hasta el primer audio muy corto mantiene ágil la conversación y preserva el ritmo natural de las experiencias de voz interactivas.

#1 en el Benchmark de latencia TTS
0ms250ms500ms750ms1000msTTFB p50TTFA p50TTFA p95Breeze TTS 2Breeze TTS 2 — TTFB p50 119.4ms, TTFA p50 133.6ms, TTFA p95 163.3msElevenLabs Flash v2.5ElevenLabs Flash v2.5 — TTFB p50 134.8ms, TTFA p50 154.5ms, TTFA p95 207.6msFish Audio S2.1 ProFish Audio S2.1 Pro — TTFB p50 173.6ms, TTFA p50 175.3ms, TTFA p95 271.4msInworld TTS-2Inworld TTS-2 — TTFB p50 163.1ms, TTFA p50 189.4ms, TTFA p95 220.4msCartesia Sonic 3.5Cartesia Sonic 3.5 — TTFB p50 106.8ms, TTFA p50 241.9ms, TTFA p95 344.2msxAI TTSxAI TTS — TTFB p50 249.1ms, TTFA p50 318.3ms, TTFA p95 361.1msSpeechify Simba 3.2Speechify Simba 3.2 — TTFB p50 374.5ms, TTFA p50 379.0ms, TTFA p95 424.5msAsync Flash 1.5Async Flash 1.5 — TTFB p50 247.0ms, TTFA p50 438.1ms, TTFA p95 479.6msElevenLabs v3ElevenLabs v3 — TTFB p50 544.0ms, TTFA p50 628.7ms, TTFA p95 917.9ms

Breeze TTS 2 ocupa el primer puesto del Benchmark de latencia TTS, con el menor tiempo hasta el primer audio tanto en p50 como en p95. Su latencia baja y constante mantiene ágiles las interacciones en tiempo real y preserva el ritmo natural de la conversación.

Streaming de Breeze TTS 2 en tiempo real

La API en tiempo real mantiene abierta una única sesión WebSocket a lo largo de todos los turnos de la conversación: agrega texto a medida que llega y el audio PCM sin procesar vuelve en streaming mientras el modelo todavía está generando.

import { BreezeBlueClient } from "@breeze.blue/sdk";

const client = new BreezeBlueClient();

// One WebSocket session, many conversation turns.
const connection = await client.textToSpeech.realtime.connect("voc_...", {
  modelId: "breeze-tts-2",
});

const consumer = (async () => {
  for await (const message of connection) {
    if (message.type === "audio") {
      play(message.audio); // raw PCM, streamed while the model generates
    } else if (message.type === "turn.done") {
      return;
    }
  }
})();

connection.startTurn("turn_1");
connection.appendText("Hello from Breeze TTS 2.");
connection.flush();
connection.endTurn();

await consumer;
connection.close();
Consigue tu API key de BreezeBlue →

CAPACIDAD 04

Voz multilingüe

Crea voces de IA de texto a voz naturales y expresivas en 50 idiomas, con control de acento para adaptar la entrega a distintas regiones y audiencias.

Elige un idioma

Every journey finds its meaning when someone dares to take the first step.

Original: Every journey finds its meaning when someone dares to take the first step.

Dale vida a tu guion en BreezeBlue Creator →

BreezeBlue Creator

Diseña, dirige y genera voces en una sola herramienta de creación que funciona en el navegador.

BreezeBlue Creator text-to-speech workspace

API y SDK de Texto a voz

Integra Texto a voz de BreezeBlue en tu producto mediante APIs o SDKs.

BreezeBlue Text to Speech API code sample

Solución empresarial

Lleva experiencias de voz de nivel producción a tu organización

Enterprise voice experience illustration