Saved in:
Bibliographic Details
Main Authors: Anand, Srija, Sankar, Ashwin, Sethi, Ishvinder, Pareek, Aaditya, Rajput, Kartik, Yadav, Gaurav, Narasimhan, Nikhil, Pandya, Adish, Halder, Deepon, Khan, Mohammed Safi Ur Rahman, S V, Praveen, Banga, Shobhit, Khapra, Mitesh M
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.21481
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917431096639488
author Anand, Srija
Sankar, Ashwin
Sethi, Ishvinder
Pareek, Aaditya
Rajput, Kartik
Yadav, Gaurav
Narasimhan, Nikhil
Pandya, Adish
Halder, Deepon
Khan, Mohammed Safi Ur Rahman
S V, Praveen
Banga, Shobhit
Khapra, Mitesh M
author_facet Anand, Srija
Sankar, Ashwin
Sethi, Ishvinder
Pareek, Aaditya
Rajput, Kartik
Yadav, Gaurav
Narasimhan, Nikhil
Pandya, Adish
Halder, Deepon
Khan, Mohammed Safi Ur Rahman
S V, Praveen
Banga, Shobhit
Khapra, Mitesh M
contents Crowdsourced pairwise evaluation has emerged as a scalable approach for assessing foundation models. However, applying it to Text to Speech(TTS) introduces high variance due to linguistic diversity and multidimensional nature of speech perception. We present a controlled multidimensional pairwise evaluation framework for multilingual TTS that combines linguistic control with perceptually grounded annotation. Using 5K+ native and code-mixed sentences across 10 Indic languages, we evaluate 7 state-of-the-art TTS systems and collect over 120K pairwise comparisons from over 1900 native raters. In addition to overall preference, raters provide judgments across 6 perceptual dimensions: intelligibility, expressiveness, voice quality, liveliness, noise, and hallucinations. Using Bradley-Terry modeling, we construct a multilingual leaderboard, interpret human preference using SHAP analysis and analyze leaderboard reliability alongside model strengths and trade-offs across perceptual dimensions.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21481
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages
Anand, Srija
Sankar, Ashwin
Sethi, Ishvinder
Pareek, Aaditya
Rajput, Kartik
Yadav, Gaurav
Narasimhan, Nikhil
Pandya, Adish
Halder, Deepon
Khan, Mohammed Safi Ur Rahman
S V, Praveen
Banga, Shobhit
Khapra, Mitesh M
Computation and Language
Crowdsourced pairwise evaluation has emerged as a scalable approach for assessing foundation models. However, applying it to Text to Speech(TTS) introduces high variance due to linguistic diversity and multidimensional nature of speech perception. We present a controlled multidimensional pairwise evaluation framework for multilingual TTS that combines linguistic control with perceptually grounded annotation. Using 5K+ native and code-mixed sentences across 10 Indic languages, we evaluate 7 state-of-the-art TTS systems and collect over 120K pairwise comparisons from over 1900 native raters. In addition to overall preference, raters provide judgments across 6 perceptual dimensions: intelligibility, expressiveness, voice quality, liveliness, noise, and hallucinations. Using Bradley-Terry modeling, we construct a multilingual leaderboard, interpret human preference using SHAP analysis and analyze leaderboard reliability alongside model strengths and trade-offs across perceptual dimensions.
title Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages
topic Computation and Language
url https://arxiv.org/abs/2604.21481