패러다임 전환
콘텐츠 제작에서 실시간 음성 AI 경험으로
생성형 AI의 첫 번째 물결은 콘텐츠 제작을 바꿔 놓았어요. 이제 텍스트와 이미지, 영상, 음성을 필요할 때 바로 만들어 낼 수 있고, 프롬프트 하나로 아이디어가 콘텐츠가 돼요. 다만 이렇게 만들어진 콘텐츠는 대부분 한 번 완성된 뒤 그대로 소비되는 결과물이에요. 다음 물결이 가져오는 것은 상호작용이에요. 실시간 음성 AI가 캐릭터와 게임, 컴패니언, 인터랙티브 스토리, 보이스 에이전트를 움직이며 사용자와 맥락, 매 순간에 끊임없이 반응해요.
이제 콘텐츠는 그것을 경험하는 사람에게 계속 반응해요.
이 패러다임 전환은 사용자가 AI 음성 합성에 기대하는 바를 자연스러움 너머로 넓혀요. 자연스러운 음성은 여전히 좋은 음성 경험의 토대이고, 여기에 상호작용하는 경험은 세 가지 요건을 더 얹어요.
다양성
모든 캐릭터에게 저마다 다른 음성을 부여해요.
제어 가능성
의도와 맥락이 달라지는 대로 매 연기를 조정해요.
낮은 지연
상호작용 본래의 자연스러운 리듬을 지켜요.
이 능력들이 함께할 때, 음성은 캐릭터에 맞춰 디자인되고 그 순간에 맞춰 연출되며 실시간으로 전달돼요.
새로운 모델
Breeze TTS 2, 상호작용 경험을 위해 만든 새로운 음성 모델.
Breeze TTS 2는 자연어 기반 음성 디자인과 어떤 음성에나 적용되는 음성 디렉션, 그리고 낮은 지연의 스트리밍 생성을 하나의 모델에 담았어요. 차세대 실시간 음성 AI를 위해 만들어져, 어떤 음성이든 디자인하고 모든 연기를 연출해 실시간으로 살아 움직이게 해요.
이런 능력을 발전시키려면 그것을 제대로 측정할 방법도 필요해요. TTS 평가는 오랫동안 자연스러움을 중심에 두어 왔고, 음성 디자인과 음성 디렉션, 지연 시간은 상대적으로 체계적인 탐구가 부족했어요. Breeze TTS 2는 이 능력들을 상호작용형 음성 경험을 위한 하나의 모델로 모아, 창작자와 개발자가 음성을 어떻게 디자인하고 연기하며 실시간으로 전달할지 더 세밀하게 제어할 수 있게 해요. 이 영역의 발전을 계속 이어가기 위해 Voice Design, Voice Direction, Latency 오픈소스 벤치마크도 함께 공개해요. 이제 Breeze TTS 2를 정의하는 네 가지 능력, 즉 음성 디자인과 음성 디렉션, 낮은 지연, 다국어 음성을 하나씩 살펴볼게요.
능력 01
음성 디자인 - 모든 캐릭터에게 알맞은 AI 음성
주인공 한 명부터 수천 명의 NPC까지, Breeze TTS 2는 모든 캐릭터에게 자기만의 음성을 줘요. 미리 준비된 음성 라이브러리에만 기대지 않고, 머릿속에 그린 캐릭터를 설명해 AI 음성 합성으로 각 역할과 성격, 세계관에 더 잘 어울리는 개성 있는 음성을 만들 수 있어요.
Breeze TTS 2
MiMo-V2.5-TTS
Inworld Voice Design
Fun-AudioGen-VD
Eleven v3
Qwen3-TTS-VD
VoxCPM2Breeze TTS 2는 TTS 음성 디자인 벤치마크에서 1위를 차지했어요. Role Fit 점수가 78.02로 가장 높고 Voice Diversity도 708로 가장 크며, Transcript Pass 비율은 98.5%를 기록했어요. Role Fit에서는 2위 모델을 5.24점 앞서고, 가장 근접한 경쟁 모델보다 39% 더 많은 서로 다른 음성을 만들어 내요.
Animation & Comics
PROMPT
Masculine, mature adult. Massive heavy-set alien. Deep, gravelly, resonant bass-baritone. Gruff, authoritative, commanding, drill-sergeant cadence. Fast, high-intensity. Rough, weathered, immense power, stern mentorship.
SCRIPT
Chest up, eyes forward. Space does not care if you're tired, scared, or homesick; it only rewards discipline. Stay on my count, recruit, and I'll drag a champion out of you yet.
Film & TV
PROMPT
Young adult female, mid-20s, American accent. Professional broadcast journalist turning frantic. Clear, articulate but strained by panic. High tension, rapid pace, breathless. Polished to raw desperation.
SCRIPT
This is Elena Vance, reporting live from the east block where the crowd just broke the barricade—if anyone can hear me, we need medical teams now, please, they're getting closer.
Games
PROMPT
Female, young adult, ethereal trickster. High-pitched, breathy, mischievous. Fast, rhythmic delivery with giggles. Mocking sing-song riddles. Playful, taunting, otherworldly.
SCRIPT
Oh, you heard the whisper too? Clever little heartbeat. Follow the violet shimmer if you dare, but mind your shadow, love. It tends to trade secrets when the veil grows thin.
Literature & Stage
PROMPT
Male, mid-20s to early 30s. Smooth velvety tenor-range speaking voice with a light rasp. Soulful and deeply emotive, shifting from weary struggle to optimistic conviction. Impeccable diction, gentle Southern lilt, deliberate pace, and broad expressive pitch arcs.
SCRIPT
I've walked through too many long nights to mistake silence for peace, so when I speak of tomorrow, believe me, brother, I mean a dawn we can finally stand in together.
Myth & Folklore
PROMPT
Mature nomadic storyteller. Weathered, resonant speaking voice with rhythmic pacing and expressive pitch contours. Wide dynamic range from hushed whispers to booming proclamations; soulful.
SCRIPT
Come closer to the fire, child; the wind is telling old truths tonight. Hear how it bends the grass, then remember this: every road you take is also taking measure of you.
Research & News
PROMPT
Mature male, 50s-60s, Japanese-accented English. Dignified, measured, stoic tone. Clear texture with slight age rasp. Deliberate pacing, significant pauses. Conveys wisdom, authority, and visionary caution.
SCRIPT
We do not rebuild a nation with noise or haste. We place each stone with care, and when doubt rises, we answer it with discipline, patience, and an unshaken sense of duty.
Web & Periodical
PROMPT
Ancient male, massive scale. Deep, resonant bass, gravelly, subterranean. Slow, deliberate, aristocratic. Terrifying power, sibilant, predatory intelligence.
SCRIPT
Careful where you place your little hands, treasure-seeker; every coin in this vault knows its master, and I am not inclined to forgive even the smallest theft.
능력 02
음성 디렉션 - 모든 순간에 알맞은 연기
자연어 지시와 인라인 태그로 감정과 의도, 속도, 전달 방식을 다듬으면서도 캐릭터의 개성은 그대로 유지해요. 덕분에 창작자와 개발자는 음성 자체를 다시 디자인하지 않고도 AI 음성이 순간마다 어떻게 연기할지 더 정교하게 제어할 수 있어요.
Breeze TTS 2
MiMo-v2.5-TTS
StepAudio 2.5 TTS
Qwen-Audio 3.0-TTS-Plus
Inworld TTS-2
VoxCPM2Breeze TTS 2는 음성 디렉션 점수 4.25로 가장 높은 점수를 받아 2위 모델을 13% 앞섰어요. 여기에 SIM 점수 0.67까지 더해 보면, 감정과 의도, 속도, 전달 방식 전반에서 자연어 디렉션을 따르면서도 선택한 음성을 서로 다른 연기 속에서 계속 알아볼 수 있게 유지한다는 뜻이에요.
Communicative Intent
PROMPT
Speak with an alluring, hushed intensity, using a hypnotic and persuasive rhythm to draw the listener in and convince them to acquire a rare item.
SCRIPT
You feel the energy radiating from it, don't you? This isn't just a piece of quartz. It's a conduit. For a small offering, you can take it home and finally clear that dark cloud hanging over your future. Trust me, you need this.
REFERENCE
OUTPUT
Composition
PROMPT
Speaking while heavily out of breath and gasping for air, trying to reassure the listener that everything is fine.
SCRIPT
[gasps] I'm okay... [pants] just give me a second to catch my breath. [gasps] I promise, I'm completely fine... just ran a little too fast.
REFERENCE
OUTPUT
Emotion
PROMPT
The speaker sounds acerbic and bitter, delivering the lines with heavy, passive-aggressive sarcasm.
SCRIPT
Enjoy your vacation while I stay here and clean up the mess you left behind for me. Have a wonderful time.
REFERENCE
OUTPUT
Role
PROMPT
Speak with a gritty, swaggering, and theatrical tone, projecting an aggressive and boastful attitude of a seafaring marauder.
SCRIPT
Drop the anchor and bring out the gold! If any of you scallywags think about crossing me, you'll be swimming with the sharks before sunset. This ocean belongs to me!
REFERENCE
OUTPUT
Variation
PROMPT
Start with a confident, reassuring tone. At the word 'Wait', suddenly shift to terrified, breathless alarm as if reacting to a frightening noise.
SCRIPT
There's absolutely nothing to worry about, I've done this a hundred times before. Wait... did you hear that sound? Something's definitely not right here.
REFERENCE
OUTPUT
능력 03
낮은 지연 - 대화의 속도로 움직이는 실시간 음성 AI
Breeze TTS 2는 자연스러움과 음성 디자인, 음성 디렉션을 실시간 음성 AI 애플리케이션을 위해 만든 하나의 저지연 모델에 담았어요. 첫 오디오까지의 시간이 짧아 대화가 즉각적으로 이어지고, 상호작용형 음성 경험 본래의 자연스러운 리듬이 유지돼요.
Breeze TTS 2는 TTS 지연 시간 벤치마크에서 1위를 차지해 p50과 p95 모두에서 가장 빠른 첫 오디오 도달 시간을 기록했어요. 꾸준히 낮은 지연 덕분에 실시간 상호작용이 즉각적으로 반응하고, 대화 본래의 자연스러운 리듬이 유지돼요.
Breeze TTS 2 실시간 스트리밍
Realtime API는 대화 턴이 이어지는 동안 하나의 WebSocket 세션을 계속 열어 둬요. 텍스트가 도착하는 대로 이어 붙이면, 모델이 아직 생성하는 중에도 원본 PCM 오디오가 스트리밍으로 돌아와요.
import { BreezeBlueClient } from "@breeze.blue/sdk";
const client = new BreezeBlueClient();
// One WebSocket session, many conversation turns.
const connection = await client.textToSpeech.realtime.connect("voc_...", {
modelId: "breeze-tts-2",
});
const consumer = (async () => {
for await (const message of connection) {
if (message.type === "audio") {
play(message.audio); // raw PCM, streamed while the model generates
} else if (message.type === "turn.done") {
return;
}
}
})();
connection.startTurn("turn_1");
connection.appendText("Hello from Breeze TTS 2.");
connection.flush();
connection.endTurn();
await consumer;
connection.close();능력 04
다국어 음성
50개 언어에 걸쳐 자연스럽고 표현력 있는 AI 음성 합성 음성을 만들고, 악센트 제어로 지역과 청중에 맞게 전달 방식을 조정해요.
언어 선택
Every journey finds its meaning when someone dares to take the first step.
원문: Every journey finds its meaning when someone dares to take the first step.



