Provider-specific pre-test protocol

ElevenLabs benchmark test protocol

This page defines the exact evidence we want before publishing an ElevenLabs benchmark score. It is a protocol, not a review result.

Current product facts to record

Do not benchmark “ElevenLabs” as one undefined configuration.

ElevenLabs currently documents multiple TTS model families and configurable voice settings. Every run should record the exact model ID, voice, plan/tier, output format and settings used. For models that expose stability, similarity, style, speed or speaker boost, record each value. If a model does not expose a setting, record that fact instead of inventing a default.

Minimum test matrix

Run at least these comparable passes

Quality-oriented pass

Use the highest-quality generally available TTS model appropriate for the tested voice and record the exact model ID.

Low-latency pass

Use the provider's low-latency model suitable for interactive or real-time workflows and record model ID plus output format.

Hindi/Hinglish pass

Run the public India benchmark pack unchanged and record the selected voice's language/accent context.

Long-form pass

Generate several minutes from a fixed passage and note drift in voice identity, pacing, pronunciation and loudness.

Settings capture

Record settings before listening

FieldRecord
Model IDExact API/UI model identifier
VoiceVoice name and voice ID when available
StabilityNumeric value or “not exposed for model”
SimilarityNumeric value or “not exposed for model”
StyleNumeric value or “not exposed for model”
Speaker boostOn / off / unavailable
SpeedNumeric value or unavailable
Output formatCodec, sample rate and bitrate where known
Plan/tierPlan used during the test
Test dateUTC date of generation
Standard scripts

Use fixed input

  1. Run the global VoicePilot AI Voice Benchmark without simplifying difficult phrases.
  2. Run the Hindi/Hinglish benchmark pack unchanged.
  3. For each pass, retain the exact source text used.
  4. If pronunciation dictionaries or custom rules are added, publish those rules with the record.
Revision test

Measure correction cost, not only first-pass quality

  1. Replace one proper noun.
  2. Replace one numeric/currency value.
  3. Replace one complete Hinglish sentence.

Record attempts, elapsed time, whether unrelated audio had to be regenerated, and whether the replacement matches adjacent timbre, pacing and loudness.

Non-determinism

One lucky generation is not enough

For any section with a surprising success or failure, repeat the generation. Record whether the outcome is consistent across runs. A benchmark should reflect reproducible production behavior rather than a cherry-picked sample.

Evidence package

Minimum evidence before score publication

  • Completed run-recorder export
  • Exact source scripts
  • Model, voice and settings
  • Revision-attempt log
  • Notes on manual edits and workarounds
  • Audio evidence where publication rights permit
Commercial separation

Affiliate links do not influence the protocol

VoicePilot may earn commission from disclosed ElevenLabs referral links on commercial pages. This benchmark protocol is independent: the same fixed scripts, scoring rubric and evidence requirements apply to competing providers.