Beyond review scores

Blind Voice Arena

Same prompt, two hidden providers, multiple raters, A/B/tie. VoicePilot avoids pretending a single editor's 4.2/5 rating is objective.

Audio samples
—
Reviewable
—
Open pairs
—
Completed pairs
—
Votes
—

Why this is stronger

Pairwise listening reduces provider-brand bias and makes disagreements visible. Tie is valid, and no winner appears until minimum-vote thresholds are reached.

Four evidence layers

  1. Generation integrity
  2. Objective checks
  3. Blind listening
  4. Revision cost and time

Current state

Loading arena backend…

Sarvam's first public Bulbul V3 run remains generation evidence only because no reviewable audio artifact was captured. That is intentional.