TTS VOICE DIRECTION BENCHMARK

How well can TTS models take direction?

An open benchmark measuring how reliably TTS models follow natural-language performance instructions while preserving speaker identity from zero-shot reference audio.

LEADERBOARD

Voice Direction Leaderboard

Voice Direction Score (VDS)

RankModel
OverallFoundationalSituationalComplexSPK_SIM
1BreezeBlue logoBreeze TTS 2
4.154.084.224.170.64
2Xiaomi logoMiMo-v2.5-TTS
3.763.813.863.620.66
3StepFun logoStepAudio 2.5 TTS
3.483.663.513.280.64
4Qwen logoQwen-Audio 3.0-TTS-Plus
3.333.883.122.980.65
5Inworld logoInworld TTS-2
3.012.863.482.700.71
6OpenBMB logoVoxCPM2
2.492.512.772.190.70

METRICS

Two measures of effective voice direction

We evaluate direction following and speaker preservation. Together, they capture how well a model changes the performance and how reliably it maintains the original voice.

Get Evaluation Suite on GitHub

Voice Direction Score

Measures how successfully the generated speech fulfills the natural-language direction. It uses a holistic five-point scale, considering fulfillment of explicit requirements, requested intensity, timing and transitions, preservation of the intended meaning, and the coherence and practical usability of the performance.

Speaker Similarity

Measures how well the generated audio preserves the reference speaker. We extract speaker embeddings and compute cosine similarity between each generated sample and its reference audio. Higher values indicate stronger preservation of the original voice.

DATASET

700 cases. 9 axes. 25 voices.

Evaluate how well TTS models follow natural-language direction across 9 axes of steerability and 25 distinct reference voices. The benchmark tests whether voice control generalizes across different speakers, instructions, and scenarios.

Open Hugging Face
700Voice direction cases
9Axes of steerability
25Distinct reference voices

Capability Groups

Foundational Speech Control

Direct control of audible speech features

Accent50Acoustic Attributes100Vocal Events50
200cases

Situational Voice Acting

Delivery shaped by emotion, state, intent, and role

Emotion150Physiological State50Communicative Intent50Role100
350cases

Complex Instruction

Combining multiple controls and changing delivery over time

Composition100Variation50
150cases

LISTEN TO THE GAP

Which system follows the voice direction better?

Voice Direction

A voice choking up, fighting against a tightening throat.

Preset Voice

Reference speaker audio

Script

It's just hard to say goodbye to this old house after living here for forty years and seeing everything change around us.

Voice A

0:000:00

Voice B

0:000:00