Breeze TTS 2 のご紹介

あらゆる声をデザインし、すべての演技をディレクションします。しかもリアルタイムで。

Breeze TTS 2 のご紹介 — 新しいフラッグシップ音声モデル

パラダイムシフト

コンテンツ制作から、リアルタイム音声 AI 体験へ

生成 AI の第一の波は、コンテンツ制作を変えました。テキスト、画像、動画、音声は、いまやプロンプトひとつで必要なときに作り出せます。とはいえ、そこで生まれるものの多くは、一度作られてから消費される完成物のままです。次に来る波がもたらすのは、相互作用です。リアルタイム音声 AI がキャラクター、ゲーム、コンパニオン、インタラクティブな物語、そして音声エージェントを動かし、ユーザー、文脈、瞬間のひとつひとつに応答し続けます。

コンテンツは、それを体験している人に絶えず応答するようになります。

このパラダイムシフトによって、音声合成による AI 音声に期待されるものは、自然さだけにとどまらなくなります。自然に聞こえる音声は今も高品質な音声体験の土台ですが、インタラクティブな体験は、そこにさらに 3 つの要件を加えます。

多様性

すべてのキャラクターに、それぞれ異なる声を与えます。

制御性

意図と文脈の変化に合わせて、一つひとつの演技を調整します。

低レイテンシ

やり取り本来の自然なリズムを保ちます。

これらの能力が組み合わさることで、声はキャラクターのためにデザインされ、その瞬間のためにディレクションされ、リアルタイムで届けられます。


新しいモデル

Breeze TTS 2 — インタラクティブな体験のために作られた、新しい音声モデルです。

Breeze TTS 2 は、自然言語による音声デザイン、どの音声にも効く音声ディレクション、そして低レイテンシのストリーミング生成を、ひとつのモデルに統合しています。次世代のリアルタイム音声 AI のために作られており、あらゆる声をデザインし、すべての演技をディレクションし、それをリアルタイムで動かせます。

こうした能力を前に進めるには、それを測る方法も良くなければなりません。TTS の評価はこれまで自然さを中心に据えてきたため、音声デザイン音声ディレクションレイテンシは体系的な検証が進んでいませんでした。Breeze TTS 2 はこれらの能力を、インタラクティブな音声体験のために作られたひとつのモデルにまとめ、音声がどのようにデザインされ、演じられ、リアルタイムで届けられるかを、クリエイターと開発者がこれまで以上に制御できるようにします。この領域の進歩を後押しするため、音声デザイン、音声ディレクション、レイテンシの各ベンチマークもオープンソースで公開します。それでは、Breeze TTS 2 を形づくる 4 つの能力 — 音声デザイン、音声ディレクション、低レイテンシ、多言語音声 — を順に見ていきます。


能力 01

音声デザイン — すべてのキャラクターに、ふさわしい AI 音声を

Breeze TTS 2 は、ひとりの主人公から数千体の NPC まで、すべてのキャラクターに固有の声を与えます。プリセットの音声ライブラリだけに頼る必要はありません。思い描いたキャラクターを言葉で説明すれば、役柄、性格、世界観によりよく合う、特徴のある音声合成による AI 音声を生成できます。

TTS 音声デザインベンチマークで第 1 位
順位モデル
Role FitVoice Diversity
1Breeze TTS 2
78.02708
2MiMo-V2.5-TTS
72.78202
3Inworld Voice Design
64.61378
4Fun-AudioGen-VD
63.82136
5Eleven v3
58.88111
6Qwen3-TTS-VD
58.4883
7VoxCPM2
55.28509

Breeze TTS 2 は TTS 音声デザインベンチマークで第 1 位となり、最高の Role Fit スコア(78.02)と最大の Voice Diversity(708)、そして 98.5% の Transcript Pass 率を記録しました。Role Fit では 2 位のモデルを 5.24 ポイント上回り、最も近い競合より 39% 多くの異なる音声を生成します。

Animation & Comics

PROMPT

Masculine, mature adult. Massive heavy-set alien. Deep, gravelly, resonant bass-baritone. Gruff, authoritative, commanding, drill-sergeant cadence. Fast, high-intensity. Rough, weathered, immense power, stern mentorship.

SCRIPT

Chest up, eyes forward. Space does not care if you're tired, scared, or homesick; it only rewards discipline. Stay on my count, recruit, and I'll drag a champion out of you yet.

0:11

Film & TV

PROMPT

Young adult female, mid-20s, American accent. Professional broadcast journalist turning frantic. Clear, articulate but strained by panic. High tension, rapid pace, breathless. Polished to raw desperation.

SCRIPT

This is Elena Vance, reporting live from the east block where the crowd just broke the barricade—if anyone can hear me, we need medical teams now, please, they're getting closer.

0:08

Games

PROMPT

Female, young adult, ethereal trickster. High-pitched, breathy, mischievous. Fast, rhythmic delivery with giggles. Mocking sing-song riddles. Playful, taunting, otherworldly.

SCRIPT

Oh, you heard the whisper too? Clever little heartbeat. Follow the violet shimmer if you dare, but mind your shadow, love. It tends to trade secrets when the veil grows thin.

0:11

Literature & Stage

PROMPT

Male, mid-20s to early 30s. Smooth velvety tenor-range speaking voice with a light rasp. Soulful and deeply emotive, shifting from weary struggle to optimistic conviction. Impeccable diction, gentle Southern lilt, deliberate pace, and broad expressive pitch arcs.

SCRIPT

I've walked through too many long nights to mistake silence for peace, so when I speak of tomorrow, believe me, brother, I mean a dawn we can finally stand in together.

0:11

Myth & Folklore

PROMPT

Mature nomadic storyteller. Weathered, resonant speaking voice with rhythmic pacing and expressive pitch contours. Wide dynamic range from hushed whispers to booming proclamations; soulful.

SCRIPT

Come closer to the fire, child; the wind is telling old truths tonight. Hear how it bends the grass, then remember this: every road you take is also taking measure of you.

0:12

Research & News

PROMPT

Mature male, 50s-60s, Japanese-accented English. Dignified, measured, stoic tone. Clear texture with slight age rasp. Deliberate pacing, significant pauses. Conveys wisdom, authority, and visionary caution.

SCRIPT

We do not rebuild a nation with noise or haste. We place each stone with care, and when doubt rises, we answer it with discipline, patience, and an unshaken sense of duty.

0:15

Web & Periodical

PROMPT

Ancient male, massive scale. Deep, resonant bass, gravelly, subterranean. Slow, deliberate, aristocratic. Terrifying power, sibilant, predatory intelligence.

SCRIPT

Careful where you place your little hands, treasure-seeker; every coin in this vault knows its master, and I am not inclined to forgive even the smallest theft.

0:16
Voice Galaxy で 15,000+ のキャラクターボイスを探す →

能力 02

音声ディレクション — すべての瞬間に、ふさわしい演技を

自然言語の指示とインラインタグで、感情、意図、テンポ、話し方を形づくりながら、どのキャラクターも聞き分けられる状態に保てます。音声そのものを作り直さなくても、AI 音声が瞬間ごとにどう演じるかを、クリエイターと開発者がこれまで以上に細かく制御できます。

TTS 音声ディレクションベンチマークで第 1 位
順位モデル
音声ディレクションスコアSPK_SIM
1Breeze TTS 2
4.250.67
2MiMo-v2.5-TTS
3.760.66
3StepAudio 2.5 TTS
3.480.64
4Qwen-Audio 3.0-TTS-Plus
3.330.65
5Inworld TTS-2
3.010.71
6VoxCPM2
2.490.70

Breeze TTS 2 は最高の音声ディレクションスコア(4.25)を達成し、2 位のモデルを 13% 上回りました。0.67 という SIM スコアと合わせて見ると、この結果は、感情、意図、テンポ、話し方にわたる自然言語のディレクションに従いながら、選んだ音声を演技が変わっても聞き分けられる状態に保てることを示しています。

Communicative Intent

PROMPT

Speak with an alluring, hushed intensity, using a hypnotic and persuasive rhythm to draw the listener in and convince them to acquire a rare item.

SCRIPT

You feel the energy radiating from it, don't you? This isn't just a piece of quartz. It's a conduit. For a small offering, you can take it home and finally clear that dark cloud hanging over your future. Trust me, you need this.

REFERENCE

0:20

OUTPUT

0:13

Composition

PROMPT

Speaking while heavily out of breath and gasping for air, trying to reassure the listener that everything is fine.

SCRIPT

[gasps] I'm okay... [pants] just give me a second to catch my breath. [gasps] I promise, I'm completely fine... just ran a little too fast.

REFERENCE

0:19

OUTPUT

0:15

Emotion

PROMPT

The speaker sounds acerbic and bitter, delivering the lines with heavy, passive-aggressive sarcasm.

SCRIPT

Enjoy your vacation while I stay here and clean up the mess you left behind for me. Have a wonderful time.

REFERENCE

0:31

OUTPUT

0:06

Role

PROMPT

Speak with a gritty, swaggering, and theatrical tone, projecting an aggressive and boastful attitude of a seafaring marauder.

SCRIPT

Drop the anchor and bring out the gold! If any of you scallywags think about crossing me, you'll be swimming with the sharks before sunset. This ocean belongs to me!

REFERENCE

0:28

OUTPUT

0:11

Variation

PROMPT

Start with a confident, reassuring tone. At the word 'Wait', suddenly shift to terrified, breathless alarm as if reacting to a frightening noise.

SCRIPT

There's absolutely nothing to worry about, I've done this a hundred times before. Wait... did you hear that sound? Something's definitely not right here.

REFERENCE

0:16

OUTPUT

0:09
BreezeBlue Creator で自分の声を無料でディレクションする →

能力 03

低レイテンシ — 会話の速度で動くリアルタイム音声 AI

Breeze TTS 2 は、自然さ、音声デザイン、音声ディレクションを、リアルタイム音声 AI アプリケーションのために作られたひとつの低レイテンシモデルにまとめています。初回音声までが速いことで、会話は応答性を保ち、インタラクティブな音声体験本来の自然なリズムが損なわれません。

TTS レイテンシベンチマークで第 1 位
0ms250ms500ms750ms1000msTTFB p50TTFA p50TTFA p95Breeze TTS 2Breeze TTS 2 — TTFB p50 119.4ms, TTFA p50 133.6ms, TTFA p95 163.3msElevenLabs Flash v2.5ElevenLabs Flash v2.5 — TTFB p50 134.8ms, TTFA p50 154.5ms, TTFA p95 207.6msFish Audio S2.1 ProFish Audio S2.1 Pro — TTFB p50 173.6ms, TTFA p50 175.3ms, TTFA p95 271.4msInworld TTS-2Inworld TTS-2 — TTFB p50 163.1ms, TTFA p50 189.4ms, TTFA p95 220.4msCartesia Sonic 3.5Cartesia Sonic 3.5 — TTFB p50 106.8ms, TTFA p50 241.9ms, TTFA p95 344.2msxAI TTSxAI TTS — TTFB p50 249.1ms, TTFA p50 318.3ms, TTFA p95 361.1msSpeechify Simba 3.2Speechify Simba 3.2 — TTFB p50 374.5ms, TTFA p50 379.0ms, TTFA p95 424.5msAsync Flash 1.5Async Flash 1.5 — TTFB p50 247.0ms, TTFA p50 438.1ms, TTFA p95 479.6msElevenLabs v3ElevenLabs v3 — TTFB p50 544.0ms, TTFA p50 628.7ms, TTFA p95 917.9ms

Breeze TTS 2 は TTS レイテンシベンチマークで第 1 位となり、p50 と p95 のどちらでも最速の初回音声までの時間を実現しています。安定して低いレイテンシが、リアルタイムのやり取りの応答性と、会話本来の自然なリズムを保ちます。

Breeze TTS 2 をリアルタイムでストリーミングする

Realtime API は、会話のターンをまたいで 1 本の WebSocket セッションを開いたままにします。テキストは届いた分から追加でき、モデルが生成を続けている間も、raw PCM 音声がストリームで返ってきます。

import { BreezeBlueClient } from "@breeze.blue/sdk";

const client = new BreezeBlueClient();

// One WebSocket session, many conversation turns.
const connection = await client.textToSpeech.realtime.connect("voc_...", {
  modelId: "breeze-tts-2",
});

const consumer = (async () => {
  for await (const message of connection) {
    if (message.type === "audio") {
      play(message.audio); // raw PCM, streamed while the model generates
    } else if (message.type === "turn.done") {
      return;
    }
  }
})();

connection.startTurn("turn_1");
connection.appendText("Hello from Breeze TTS 2.");
connection.flush();
connection.endTurn();

await consumer;
connection.close();
BreezeBlue の API key を取得する →

能力 04

多言語音声

50 言語にわたって、自然で表現力豊かな音声合成による AI 音声を作成できます。アクセント制御によって、地域やオーディエンスに合わせて話し方を調整できます。

言語を選択

Every journey finds its meaning when someone dares to take the first step.

原文: Every journey finds its meaning when someone dares to take the first step.

BreezeBlue Creator で台本に命を吹き込む →

BreezeBlue Creator

ブラウザだけで完結するクリエイターツールで、音声のデザイン、ディレクション、生成ができます。

BreezeBlue Creator text-to-speech workspace

音声合成 API と SDK

BreezeBlue の音声合成を、API または SDK で製品に組み込めます。

BreezeBlue Text to Speech API code sample

エンタープライズソリューション

本番品質の音声体験を、あなたの組織へ

Enterprise voice experience illustration