Research infrastructure

VoicePilot TTS Benchmark Index

A public, production-oriented framework for comparing text-to-speech platforms with the same script, the same revision tasks and the same scoring rules. No vendor receives a score until a reproducible test record exists.

Status

Leaderboard data collection is open.

We are building the first VoicePilot production benchmark dataset. Candidate platforms are listed below for coverage planning only. A blank score means not yet independently tested — it is not a negative result.

Candidate provider roster

Platforms queued for comparable testing

PlatformEnglishHindi/HinglishRevision testLong-formStatus
ElevenLabs————Queued
Murf————Queued
PlayHT————Queued
Speechify Studio————Queued
Google Cloud Text-to-Speech————Queued
Microsoft Azure AI Speech————Queued
Amazon Polly————Queued
Cartesia————Queued
OpenAI text-to-speech————Queued
Deepgram Aura————Queued
Scoring model

Six production dimensions, each scored 1–5

Pronunciation

Names, technical terms, acronyms and recurring vocabulary.

Numbers & symbols

Dates, currency, percentages, times, units and model numbers.

Multilingual handling

Hindi/Hinglish switching, borrowed terms and accent stability.

Pacing & emphasis

Pauses, stress, requested delivery changes and sentence meaning.

Revision speed

Time and attempts required for word, sentence and paragraph corrections.

Long-form consistency

Voice identity, pacing and energy stability over longer audio.

Evidence standard

What is required before a score appears

  1. The same benchmark script is used without simplifying difficult phrases for a provider.
  2. The reviewer records the test date, voice/model, settings and relevant plan.
  3. Each of the three standardized revision tasks is completed.
  4. Scores are recorded in the public scorecard format.
  5. Any provider-specific workaround is disclosed in the notes.
India benchmark layer

Designed to expand beyond English-first testing

The benchmark roadmap prioritizes Hindi and Hinglish first, followed by additional Indian-language test packs. The goal is not to reward a model for a polished English demo while ignoring the multilingual production conditions creators and localization teams actually face.

For vendors & researchers

Want your platform included?

Platforms can provide access or nominate a model for testing. Participation does not guarantee a favorable score, ranking or editorial recommendation. Paid placement does not change benchmark results.