Salvato in:
| Autori principali: | Wei, Jui-Chiang, Lin, Yi-Cheng, Ritter-Gutierrez, Fabian, Lee, Hung-yi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2506.07237 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Distilling a speech and music encoder with task arithmetic
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
Dataset-Distillation Generative Model for Speech Emotion Recognition
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2024)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2024)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
SPAR-K: Scheduled Periodic Alternating Early Exit for Spoken Language Models
di: Huang, Hsiao-Ying, et al.
Pubblicazione: (2026)
di: Huang, Hsiao-Ying, et al.
Pubblicazione: (2026)
Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
di: Lu, Ke-Han, et al.
Pubblicazione: (2025)
di: Lu, Ke-Han, et al.
Pubblicazione: (2025)
MMMOS: Multi-domain Multi-axis Audio Quality Assessment
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
Gender Bias in Instruction-Guided Speech Synthesis Models
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2025)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2025)
CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
di: Tsai, Yun-Shao, et al.
Pubblicazione: (2025)
di: Tsai, Yun-Shao, et al.
Pubblicazione: (2025)
A correlation-permutation approach for speech-music encoders model merging
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2026)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2026)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2026)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2026)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
HighRateMOS: Sampling-Rate Aware Modeling for Speech Quality Assessment
di: Ren, Wenze, et al.
Pubblicazione: (2025)
di: Ren, Wenze, et al.
Pubblicazione: (2025)
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
di: Hsu, Tzu-wen, et al.
Pubblicazione: (2025)
di: Hsu, Tzu-wen, et al.
Pubblicazione: (2025)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2024)
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2024)
Meta-PerSER: Few-Shot Listener Personalized Speech Emotion Recognition via Meta-learning
di: Shen, Liang-Yeh, et al.
Pubblicazione: (2025)
di: Shen, Liang-Yeh, et al.
Pubblicazione: (2025)
How Contrastive Decoding Enhances Large Audio Language Models?
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2026)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2026)
MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
di: Huang, Hsiao-Ying, et al.
Pubblicazione: (2025)
di: Huang, Hsiao-Ying, et al.
Pubblicazione: (2025)
MOS-Bias: From Hidden Gender Bias to Gender-Aware Speech Quality Assessment
di: Ren, Wenze, et al.
Pubblicazione: (2026)
di: Ren, Wenze, et al.
Pubblicazione: (2026)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
di: Liu, Andy T., et al.
Pubblicazione: (2024)
di: Liu, Andy T., et al.
Pubblicazione: (2024)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
di: Huang, Chien-yu, et al.
Pubblicazione: (2024)
di: Huang, Chien-yu, et al.
Pubblicazione: (2024)
Parallel Synthesis for Autoregressive Speech Generation
di: Hsu, Po-chun, et al.
Pubblicazione: (2022)
di: Hsu, Po-chun, et al.
Pubblicazione: (2022)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
di: Lin, Guan-Ting, et al.
Pubblicazione: (2025)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2025)
USAD: Universal Speech and Audio Representation via Distillation
di: Chang, Heng-Jui, et al.
Pubblicazione: (2025)
di: Chang, Heng-Jui, et al.
Pubblicazione: (2025)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Distilling a speech and music encoder with task arithmetic
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025) -
ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025) -
Dataset-Distillation Generative Model for Speech Emotion Recognition
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2024) -
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024) -
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)