Describe the Character. Hear Their Voice.

Voice design turns a written description into an original, reusable voice. No recordings, no reference speaker — write who the character is, audition the result, and save it to use across Breeze Blue.

You describe

A weathered sea captain in her sixties. Low, gravelly voice, slow deliberate pacing, a faint Irish accent, warm underneath the saltiness.

You hear

A low, rough-edged female voice with unhurried delivery and a soft Irish lilt — stern on the surface, kind underneath.

Performing the line

“Tie that line down, lad — the sea doesn’t forgive twice.”

You describe

A hyperactive robot sidekick from a kids’ cartoon. High-pitched with a metallic ring, talks fast, always on the edge of giggling.

You hear

A bright, springy synthetic voice with quick pacing and giggly energy that stays perfectly clear at speed.

Performing the line

“New plan! Same as the old plan, but with way more jetpacks!”

You describe

A calm meditation guide, androgynous, mid-thirties. Soft and breathy, close to the microphone, almost a whisper but perfectly clear.

You hear

An intimate, airy voice with gentle pacing and warm low mids — quiet without ever losing intelligibility.

Performing the line

“Let the thought pass by, like a cloud crossing a wide sky.”

From written description to saved voice in three steps

The whole loop — describe, audition, save — happens in one workspace, and the voice you keep becomes a permanent asset.

1

Describe

Write who the character is in plain language — age, texture, pacing, accent, attitude. If the character is already drawn, upload the artwork and start from the image instead.

2

Audition

Generate a candidate voice and hear it perform a preview line. Not quite right? Adjust the description — deeper, slower, less polished — and generate again until the voice matches the one in your head.

3

Save

Save the keeper to your voice list. From that moment it is a reusable asset: pick it in Text to Speech, cast it in Studio projects, and reference it by ID from the API.

How to write a description that works

Concrete, audible traits beat abstract personality words. The strongest descriptions state a few of these directly:

Age & gender
“a woman in her sixties”
Texture
“gravelly”, “breathy”, “metallic”
Pacing
“slow and deliberate”, “talks fast”
Accent
“a faint Irish accent”
Role & attitude
“a tired night-shift detective”

One or two sentences is enough. Keep what the voice sounds like; drop what the character merely does in the story.

When voice design is the right tool

Five situations creators bring to voice design — and exactly what happens in each one.

01

The preset voice is always almost right

You can hear the voice in your head, and every ready-made option lands slightly off — the right tone but the wrong age, the right accent but too polished. Voice design closes that last gap: write down exactly the voice you are hearing, audition the result, and sharpen the description until it clicks. And when close-enough genuinely is enough, starting from a curated voice stays the faster route.

Pick a starting point in the Voice Library
02

The character has no reference speaker

Game NPCs, story characters, virtual hosts — original characters do not come with a recording you could clone. Voice design needs no source audio at all: the description is the source. Turn the character sheet into a sentence or two, and the result is an original voice with no human speaker behind it.

See how creators build full character casts
03

The character is drawn but not yet heard

Sometimes the art arrives before the voice. In the voice design workspace you can start from the character image itself: upload the artwork and generate a voice shaped to match what you see, adding a written description alongside to steer what the image cannot show — pacing, accent, attitude.

Start a voice from character art
04

Shortlist up to 3 candidates per API request

On the web you audition one candidate at a time and refine the description between takes. The Voice Design API works in bigger strokes: one request to the design endpoint returns up to 3 candidate voices from the same description, so a pipeline can generate a shortlist programmatically and let a human pick the keeper.

Read the Voice Design API docs
POST /v1/text-to-voice/design
{
  "voice_description": "A weathered sea captain in her
    sixties, low and gravelly, with a faint Irish accent.",
  "text": "Tie that line down, lad.",
  "preview_count": 3
}
05

You love the voice — you need new performances

Then you are done designing. When the same voice should whisper one line and shout the next, keep the voice and direct the delivery instead: Breeze TTS 2 takes plain-language instructions line by line. Voice design decides who speaks; voice direction decides how each line is performed.

Direct performances in text to speech

Design, Library, Clone, or Direction?

Four tools, four starting points. Choose by what you already have.

Voice Library

Start from a finished voice. Browse curated voices by style and language, preview them instantly, and use them as-is — the fastest route when a ready-made voice fits.

Explore the Voice Library

Voice Cloning

Start from a real recording. When the voice already exists and you have the speaker’s permission, cloning reproduces that specific voice from a short authorized sample.

How 5-second voice cloning works

Voice Design

Start from a description — that is this page. When the voice exists only in your head or in character art, design creates an original voice with no recording and no reference speaker.

Voice Direction

Start from a voice you already trust. Direction changes how a line is delivered — emotion, pacing, delivery style — without changing who is speaking.

Direct any voice with plain-language instructions

Evidence, not adjectives

The open Voice Design Benchmark scores 1,000 character voice-design tasks across leading models (evaluated August 2026). Breeze TTS 2 ranks first on role fit and voice diversity:

78.02

Role Fit

How closely generated voices match the described role, normalized to 100 — the top score among evaluated models.

708

Distinct voice clusters

Distinct WavLM speaker clusters across the full task set — the widest voice diversity in the evaluation.

98.5%

Transcript pass

Share of generated clips that clearly speak their intended preview line.

See the full Voice Design Benchmark leaderboard and methodology

Voice Design Benchmark · evaluated August 2026

From the browser to the API

Design in the web app

Describe the character — or upload their art — audition candidates, and save the keeper. Saved voices join your voice list immediately, ready to pick in Text to Speech and to cast alongside other characters in Studio projects.

Scale with the Voice Design API

Generate candidates with POST /v1/text-to-voice/design, then persist the one you keep with POST /v1/text-to-voice. Saved voices are addressable by ID from the same endpoints your text to speech integration already uses.

Up to 3 candidates per design request through the API; the web workspace auditions one candidate at a time.

Frequently asked questions

What is AI voice design?

Voice design creates an original voice from a written description or a character image instead of a recording. You describe the speaker — age, texture, pacing, accent, attitude — audition the generated voice, and save it as a reusable asset for text to speech, Studio projects, and the API.

Do I need a voice sample or a recording?

No. Voice design needs no source audio and no reference speaker — the description alone is the input. If you do have an authorized recording of the exact voice you want, voice cloning is the better tool for that job.

Whose voice is a designed voice?

No one’s. A designed voice is generated from your description; it is not a recording of a person and is not built to reproduce any specific speaker. Similar descriptions can lead to similar-sounding results, so treat a designed voice as an original creation rather than a guaranteed one-of-a-kind sound.

Can I get several candidate voices from one description?

In the web app you audition one candidate at a time and refine the description between attempts. Through the API, a single request to the design endpoint can return up to 3 candidates, which suits pipelines that shortlist programmatically before a person picks the final voice.

Where can I use a designed voice after saving it?

Everywhere in Breeze Blue. Once saved, the voice appears in your voice list: pick it in Text to Speech, cast it in multi-character Studio projects, and reference its ID in API calls — same voice, same character, in every surface.

Can the same designed voice deliver different emotions?

Yes — but that is voice direction, not voice design. Keep the saved voice and give per-line instructions in text to speech: whisper this line, build excitement on the next. Design fixes who speaks; direction changes how each line is performed.

How much does voice design cost?

Voice design runs on the same credit balance as the rest of Breeze Blue, and you can start on the free tier. See the pricing page for current plans and how credits map to generation.