On the social bias of speech self-supervised models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lin, Yi-Cheng, Lin, Tzu-Quan, Lin, Hsi-Che, Liu, Andy T., Lee, Hung-yi |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
par: Lin, Tzu-Quan, et autres
Publié: (2024)
par: Lin, Tzu-Quan, et autres
Publié: (2024)
MelHuBERT: A simplified HuBERT on Mel spectrograms
par: Lin, Tzu-Quan, et autres
Publié: (2022)
par: Lin, Tzu-Quan, et autres
Publié: (2022)
Property Neurons in Self-Supervised Speech Transformers
par: Lin, Tzu-Quan, et autres
Publié: (2024)
par: Lin, Tzu-Quan, et autres
Publié: (2024)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
par: Liu, Andy T., et autres
Publié: (2024)
par: Liu, Andy T., et autres
Publié: (2024)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
par: Lin, Tzu-Quan, et autres
Publié: (2025)
par: Lin, Tzu-Quan, et autres
Publié: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
par: Lin, Hsi-Che, et autres
Publié: (2024)
par: Lin, Hsi-Che, et autres
Publié: (2024)
Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
par: Lin, Yi-Cheng, et autres
Publié: (2024)
par: Lin, Yi-Cheng, et autres
Publié: (2024)
How Contrastive Decoding Enhances Large Audio Language Models?
par: Lin, Tzu-Quan, et autres
Publié: (2026)
par: Lin, Tzu-Quan, et autres
Publié: (2026)
Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
par: Lin, Tsung-En, et autres
Publié: (2025)
par: Lin, Tsung-En, et autres
Publié: (2025)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
par: Lin, Tzu-Quan, et autres
Publié: (2022)
par: Lin, Tzu-Quan, et autres
Publié: (2022)
Gender Bias in Instruction-Guided Speech Synthesis Models
par: Kuan, Chun-Yi, et autres
Publié: (2025)
par: Kuan, Chun-Yi, et autres
Publié: (2025)
Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
par: Maharana, Sarthak Kumar, et autres
Publié: (2023)
par: Maharana, Sarthak Kumar, et autres
Publié: (2023)
Improving the Adversarial Robustness for Speaker Verification by Self-Supervised Learning
par: Wu, Haibin, et autres
Publié: (2021)
par: Wu, Haibin, et autres
Publié: (2021)
Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
par: Lin, Yi-Cheng, et autres
Publié: (2025)
par: Lin, Yi-Cheng, et autres
Publié: (2025)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
par: Fujita, Kenichi, et autres
Publié: (2024)
par: Fujita, Kenichi, et autres
Publié: (2024)
Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models
par: Huang, Wei-Ping, et autres
Publié: (2026)
par: Huang, Wei-Ping, et autres
Publié: (2026)
Integrating Self-supervised Speech Model with Pseudo Word-level Targets from Visually-grounded Speech Model
par: Fang, Hung-Chieh, et autres
Publié: (2024)
par: Fang, Hung-Chieh, et autres
Publié: (2024)
Multi-Distillation from Speech and Music Representation Models
par: Wei, Jui-Chiang, et autres
Publié: (2025)
par: Wei, Jui-Chiang, et autres
Publié: (2025)
A low latency attention module for streaming self-supervised speech representation learning
par: Ma, Jianbo, et autres
Publié: (2023)
par: Ma, Jianbo, et autres
Publié: (2023)
Distilling a speech and music encoder with task arithmetic
par: Ritter-Gutierrez, Fabian, et autres
Publié: (2025)
par: Ritter-Gutierrez, Fabian, et autres
Publié: (2025)
Towards Generalized Source Tracing for Codec-Based Deepfake Speech
par: Chen, Xuanjun, et autres
Publié: (2025)
par: Chen, Xuanjun, et autres
Publié: (2025)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
par: Lin, Yi-Cheng, et autres
Publié: (2024)
par: Lin, Yi-Cheng, et autres
Publié: (2024)
A correlation-permutation approach for speech-music encoders model merging
par: Ritter-Gutierrez, Fabian, et autres
Publié: (2025)
par: Ritter-Gutierrez, Fabian, et autres
Publié: (2025)
Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
par: Kuan, Chun-Yi, et autres
Publié: (2024)
par: Kuan, Chun-Yi, et autres
Publié: (2024)
Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition
par: Lin, Yi-Cheng, et autres
Publié: (2025)
par: Lin, Yi-Cheng, et autres
Publié: (2025)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
par: Lin, Yi-Cheng, et autres
Publié: (2025)
par: Lin, Yi-Cheng, et autres
Publié: (2025)
Zipformer: A faster and better encoder for automatic speech recognition
par: Yao, Zengwei, et autres
Publié: (2023)
par: Yao, Zengwei, et autres
Publié: (2023)
AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering
par: Kuan, Chun-Yi, et autres
Publié: (2026)
par: Kuan, Chun-Yi, et autres
Publié: (2026)
From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data
par: Kuan, Chun-Yi, et autres
Publié: (2025)
par: Kuan, Chun-Yi, et autres
Publié: (2025)
Tempo estimation as fully self-supervised binary classification
par: Henkel, Florian, et autres
Publié: (2024)
par: Henkel, Florian, et autres
Publié: (2024)
CR-CTC: Consistency regularization on CTC for improved speech recognition
par: Yao, Zengwei, et autres
Publié: (2024)
par: Yao, Zengwei, et autres
Publié: (2024)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
par: Gowda, Harshavardhana T., et autres
Publié: (2025)
par: Gowda, Harshavardhana T., et autres
Publié: (2025)
MMMOS: Multi-domain Multi-axis Audio Quality Assessment
par: Lin, Yi-Cheng, et autres
Publié: (2025)
par: Lin, Yi-Cheng, et autres
Publié: (2025)
Exploring speech style spaces with language models: Emotional TTS without emotion labels
par: Chandra, Shreeram Suresh, et autres
Publié: (2024)
par: Chandra, Shreeram Suresh, et autres
Publié: (2024)
MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
par: Huang, Hsiao-Ying, et autres
Publié: (2025)
par: Huang, Hsiao-Ying, et autres
Publié: (2025)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
par: Maiti, Soumi, et autres
Publié: (2023)
par: Maiti, Soumi, et autres
Publié: (2023)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
par: Wang, Hsuan-Fu, et autres
Publié: (2024)
par: Wang, Hsuan-Fu, et autres
Publié: (2024)
Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations
par: Lin, Guan-Ting, et autres
Publié: (2024)
par: Lin, Guan-Ting, et autres
Publié: (2024)
CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
par: Tsai, Yun-Shao, et autres
Publié: (2025)
par: Tsai, Yun-Shao, et autres
Publié: (2025)
TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild
par: Chang, Kai-Wei, et autres
Publié: (2026)
par: Chang, Kai-Wei, et autres
Publié: (2026)
Documents similaires
-
DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
par: Lin, Tzu-Quan, et autres
Publié: (2024) -
MelHuBERT: A simplified HuBERT on Mel spectrograms
par: Lin, Tzu-Quan, et autres
Publié: (2022) -
Property Neurons in Self-Supervised Speech Transformers
par: Lin, Tzu-Quan, et autres
Publié: (2024) -
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
par: Liu, Andy T., et autres
Publié: (2024) -
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
par: Lin, Tzu-Quan, et autres
Publié: (2025)