Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Yi-Cheng, Tsai, Yun-Shao, Chen, Kuan-Yu, Huang, Hsiao-Ying, Chou, Huang-Cheng, Lee, Hung-yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
von: Tsai, Yun-Shao, et al.
Veröffentlicht: (2025)
von: Tsai, Yun-Shao, et al.
Veröffentlicht: (2025)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)
MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
von: Huang, Hsiao-Ying, et al.
Veröffentlicht: (2025)
von: Huang, Hsiao-Ying, et al.
Veröffentlicht: (2025)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
von: Huang, Chien-yu, et al.
Veröffentlicht: (2023)
von: Huang, Chien-yu, et al.
Veröffentlicht: (2023)
Dataset-Distillation Generative Model for Speech Emotion Recognition
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
Parallel Synthesis for Autoregressive Speech Generation
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
von: Derington, Anna, et al.
Veröffentlicht: (2023)
von: Derington, Anna, et al.
Veröffentlicht: (2023)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
Toward Fairness in Speech Recognition: Discovery and mitigation of performance disparities
von: Dheram, Pranav, et al.
Veröffentlicht: (2022)
von: Dheram, Pranav, et al.
Veröffentlicht: (2022)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2024)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
How Contrastive Decoding Enhances Large Audio Language Models?
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2026)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2026)
Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
Examining the Interplay Between Privacy and Fairness for Speech Processing: A Review and Perspective
von: Leschanowsky, Anna, et al.
Veröffentlicht: (2024)
von: Leschanowsky, Anna, et al.
Veröffentlicht: (2024)
ConSep: a Noise- and Reverberation-Robust Speech Separation Framework by Magnitude Conditioning
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2024)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2024)
Towards audio language modeling -- an overview
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2025)
von: Chao, Rong, et al.
Veröffentlicht: (2025)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
von: Pham, Lam, et al.
Veröffentlicht: (2024)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
What do neural networks listen to? Exploring the crucial bands in Speech Enhancement using Sinc-convolution
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024) -
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025) -
CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
von: Tsai, Yun-Shao, et al.
Veröffentlicht: (2025) -
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024) -
VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)