Voice Reliability Index
A model can sound impressive once and still be unreliable in production. VoicePilot measures whether the same expected result appears repeatedly without retries, spelling hacks or manual correction.
Provider reliability
No provider receives a Reliability Index until at least five verified runs exist under equivalent benchmark conditions.
Why reliability is different from quality
Quality asks whether a sample sounds good. Reliability asks whether a team can expect the same acceptable result repeatedly. For production systems, a slightly less beautiful voice that succeeds on the first attempt can be more valuable than a brilliant demo that needs constant retries.
This is especially important for Indian names, rupee amounts, dates and Hinglish code-switching where small failures create disproportionate editing cost.
What counts as a failure
A run can fail because generation breaks, the target is materially wrong, a required correction is needed, a spelling workaround is required, or the result materially drifts from equivalent earlier runs.
Unknown observations do not become failures. Only verified observations enter the index.
Minimum evidence gate
The index remains blank below five verified runs. This prevents one or two lucky generations from creating a misleading reliability score. As the dataset grows, VoicePilot can later move to 10-run and 30-run confidence tiers.