Saved in:
Bibliographic Details
Main Authors: Anand, Srija, Sankar, Ashwin, Sethi, Ishvinder, Pareek, Aaditya, Rajput, Kartik, Yadav, Gaurav, Narasimhan, Nikhil, Pandya, Adish, Halder, Deepon, Khan, Mohammed Safi Ur Rahman, S V, Praveen, Banga, Shobhit, Khapra, Mitesh M
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.21481
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Crowdsourced pairwise evaluation has emerged as a scalable approach for assessing foundation models. However, applying it to Text to Speech(TTS) introduces high variance due to linguistic diversity and multidimensional nature of speech perception. We present a controlled multidimensional pairwise evaluation framework for multilingual TTS that combines linguistic control with perceptually grounded annotation. Using 5K+ native and code-mixed sentences across 10 Indic languages, we evaluate 7 state-of-the-art TTS systems and collect over 120K pairwise comparisons from over 1900 native raters. In addition to overall preference, raters provide judgments across 6 perceptual dimensions: intelligibility, expressiveness, voice quality, liveliness, noise, and hallucinations. Using Bradley-Terry modeling, we construct a multilingual leaderboard, interpret human preference using SHAP analysis and analyze leaderboard reliability alongside model strengths and trade-offs across perceptual dimensions.