The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Satish, Shree Harsha Bokkahalli, Minixhofer, Christoph, Teleki, Maria, Caverlee, James, Klejch, Ondřej, Bell, Peter, Henter, Gustav Eje, Székely, Éva |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
von: Lameris, Harm, et al.
Veröffentlicht: (2025)
von: Lameris, Harm, et al.
Veröffentlicht: (2025)
What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
von: Minixhofer, Christoph, et al.
Veröffentlicht: (2025)
von: Minixhofer, Christoph, et al.
Veröffentlicht: (2025)
TTSDS -- Text-to-Speech Distribution Score
von: Minixhofer, Christoph, et al.
Veröffentlicht: (2024)
von: Minixhofer, Christoph, et al.
Veröffentlicht: (2024)
Beyond Oversmoothing: Evaluating DDPM and MSE for Scalable Speech Synthesis in ASR
von: Minixhofer, Christoph, et al.
Veröffentlicht: (2024)
von: Minixhofer, Christoph, et al.
Veröffentlicht: (2024)
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
von: Teleki, Maria, et al.
Veröffentlicht: (2025)
von: Teleki, Maria, et al.
Veröffentlicht: (2025)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
Voice Conversion-based Privacy through Adversarial Information Hiding
von: Webber, Jacob J, et al.
Veröffentlicht: (2024)
von: Webber, Jacob J, et al.
Veröffentlicht: (2024)
Matcha-TTS: A fast TTS architecture with conditional flow matching
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
VoXtream2: Full-stream TTS with dynamic speaking rate control
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
A Practitioner's Guide to Building ASR Models for Low-Resource Languages: A Case Study on Scottish Gaelic
von: Klejch, Ondřej, et al.
Veröffentlicht: (2025)
von: Klejch, Ondřej, et al.
Veröffentlicht: (2025)
Unified speech and gesture synthesis using flow matching
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
von: Guichoux, Téo, et al.
Veröffentlicht: (2025)
von: Guichoux, Téo, et al.
Veröffentlicht: (2025)
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model
von: Wang, Siyang, et al.
Veröffentlicht: (2024)
von: Wang, Siyang, et al.
Veröffentlicht: (2024)
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Streaming Speech-to-Text Translation with a SpeechLLM
von: Parcollet, Titouan, et al.
Veröffentlicht: (2026)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2026)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
von: Jacobs, Christiaan, et al.
Veröffentlicht: (2025)
von: Jacobs, Christiaan, et al.
Veröffentlicht: (2025)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
Exploring Internal Numeracy in Language Models: A Case Study on ALBERT
von: Wennberg, Ulme, et al.
Veröffentlicht: (2024)
von: Wennberg, Ulme, et al.
Veröffentlicht: (2024)
Are These Even Words? Quantifying the Gibberishness of Generative Speech Models
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2025)
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2025)
Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026)
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026)
Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction
von: Novitasari, Sashi, et al.
Veröffentlicht: (2026)
von: Novitasari, Sashi, et al.
Veröffentlicht: (2026)
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
von: Wallbridge, Sarenne, et al.
Veröffentlicht: (2025)
von: Wallbridge, Sarenne, et al.
Veröffentlicht: (2025)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
von: Prakash, Jeena, et al.
Veröffentlicht: (2025)
von: Prakash, Jeena, et al.
Veröffentlicht: (2025)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Can LLMs Help Localize Fake Words in Partially Fake Speech?
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
Speech to Speech Synthesis for Voice Impersonation
von: Johnson, Bjorn, et al.
Veröffentlicht: (2026)
von: Johnson, Bjorn, et al.
Veröffentlicht: (2026)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
von: Storey, Edward, et al.
Veröffentlicht: (2025)
von: Storey, Edward, et al.
Veröffentlicht: (2025)
Robust Speech Activity Detection in the Presence of Singing Voice
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
von: Byun, Kyungguen, et al.
Veröffentlicht: (2025)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025) -
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025) -
Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025) -
From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026) -
Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
von: Lameris, Harm, et al.
Veröffentlicht: (2025)