Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Sheng-Lun, Liao, Yu-Ling, Chang, Yen-Hua, Huang, Hen-Hsen, Chen, Hsin-Hsi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
von: Billa, Jayadev
Veröffentlicht: (2026)
von: Billa, Jayadev
Veröffentlicht: (2026)
Towards Generalizability to Tone and Content Variations in the Transcription of Amplifier Rendered Electric Guitar Audio
von: Chen, Yu-Hua, et al.
Veröffentlicht: (2025)
von: Chen, Yu-Hua, et al.
Veröffentlicht: (2025)
ALDAS: Audio-Linguistic Data Augmentation for Spoofed Audio Detection
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024)
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024)
Listening for Expert Identified Linguistic Features: Assessment of Audio Deepfake Discernment among Undergraduate Students
von: Bhalli, Noshaba N., et al.
Veröffentlicht: (2024)
von: Bhalli, Noshaba N., et al.
Veröffentlicht: (2024)
Do Audio-Language Models Understand Linguistic Variations?
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)
Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
What Do Neurons Listen To? A Neuron-level Dissection of a General-purpose Audio Model
von: Kawamura, Takao, et al.
Veröffentlicht: (2026)
von: Kawamura, Takao, et al.
Veröffentlicht: (2026)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
Improving Audio Question Answering with Variational Inference
von: Chen, Haolin
Veröffentlicht: (2026)
von: Chen, Haolin
Veröffentlicht: (2026)
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
Joint Minimum Processing Beamforming and Near-end Listening Enhancement
von: Fuglsig, Andreas J., et al.
Veröffentlicht: (2023)
von: Fuglsig, Andreas J., et al.
Veröffentlicht: (2023)
Unsupervised Single-Channel Audio Separation with Diffusion Source Priors
von: Shi, Runwu, et al.
Veröffentlicht: (2025)
von: Shi, Runwu, et al.
Veröffentlicht: (2025)
Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2025)
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2025)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
Listen, Think, and Understand
von: Gong, Yuan, et al.
Veröffentlicht: (2023)
von: Gong, Yuan, et al.
Veröffentlicht: (2023)
Listen, Analyze, and Adapt to Learn New Attacks: An Exemplar-Free Class Incremental Learning Method for Audio Deepfake Source Tracing
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
An Empirical Analysis of Task-Induced Encoder Bias in Fréchet Audio Distance
von: Jeong, Wonwoo
Veröffentlicht: (2026)
von: Jeong, Wonwoo
Veröffentlicht: (2026)
Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis
von: Geng, Haopeng, et al.
Veröffentlicht: (2026)
von: Geng, Haopeng, et al.
Veröffentlicht: (2026)
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
PyNeuralFx: A Python Package for Neural Audio Effect Modeling
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2024)
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2024)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
Hyper Recurrent Neural Network: Condition Mechanisms for Black-box Audio Effect Modeling
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2024)
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2024)
The Voice of Equity: A Systematic Evaluation of Bias Mitigation Techniques for Speech-Based Cognitive Impairment Detection Across Architectures and Demographics
von: Haghbin, Yasaman, et al.
Veröffentlicht: (2026)
von: Haghbin, Yasaman, et al.
Veröffentlicht: (2026)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
Listenable Maps for Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
Pitch Contour Exploration Across Audio Domains: A Vision-Based Transfer Learning Approach
von: Abeßer, Jakob, et al.
Veröffentlicht: (2025)
von: Abeßer, Jakob, et al.
Veröffentlicht: (2025)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
Evaluating Speech Enhancement Systems Through Listening Effort
von: Gelderblom, Femke B., et al.
Veröffentlicht: (2024)
von: Gelderblom, Femke B., et al.
Veröffentlicht: (2024)
RF-GML: Reference-Free Generative Machine Listener
von: Biswas, Arijit, et al.
Veröffentlicht: (2024)
von: Biswas, Arijit, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
von: Billa, Jayadev
Veröffentlicht: (2026) -
Towards Generalizability to Tone and Content Variations in the Transcription of Amplifier Rendered Electric Guitar Audio
von: Chen, Yu-Hua, et al.
Veröffentlicht: (2025) -
ALDAS: Audio-Linguistic Data Augmentation for Spoofed Audio Detection
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024) -
Listening for Expert Identified Linguistic Features: Assessment of Audio Deepfake Discernment among Undergraduate Students
von: Bhalli, Noshaba N., et al.
Veröffentlicht: (2024) -
Do Audio-Language Models Understand Linguistic Variations?
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)