An Exploration of Mamba for Speech Self-Supervised Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Tzu-Quan, Kuo, Heng-Cheng, Wei, Tzu-Chieh, Cheng, Hsi-Chun, Chen, Chun Wei, Hsiao, Hsien-Fu, Tsao, Yu, Lee, Hung-yi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
Property Neurons in Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
How Contrastive Decoding Enhances Large Audio Language Models?
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2026)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2026)
Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
MelHuBERT: A simplified HuBERT on Mel spectrograms
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)
When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models
di: Li, Chen-An, et al.
Pubblicazione: (2025)
di: Li, Chen-An, et al.
Pubblicazione: (2025)
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation
di: Tsai, Yun-Shao, et al.
Pubblicazione: (2026)
di: Tsai, Yun-Shao, et al.
Pubblicazione: (2026)
Editing the Mind of Giants: An In-Depth Exploration of Pitfalls of Knowledge Editing in Large Language Models
di: Hsueh, Cheng-Hsun, et al.
Pubblicazione: (2024)
di: Hsueh, Cheng-Hsun, et al.
Pubblicazione: (2024)
Gender Bias in Instruction-Guided Speech Synthesis Models
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2025)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
di: Yang, Chih-Kai, et al.
Pubblicazione: (2023)
di: Yang, Chih-Kai, et al.
Pubblicazione: (2023)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
di: Liu, Andy T., et al.
Pubblicazione: (2024)
di: Liu, Andy T., et al.
Pubblicazione: (2024)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
di: Kuo, Tzu-Lin, et al.
Pubblicazione: (2024)
di: Kuo, Tzu-Lin, et al.
Pubblicazione: (2024)
Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
Creativity in LLM-based Multi-Agent Systems: A Survey
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging
di: Lin, Tzu-Han, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Han, et al.
Pubblicazione: (2024)
Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
di: Lu, Ke-Han, et al.
Pubblicazione: (2025)
di: Lu, Ke-Han, et al.
Pubblicazione: (2025)
Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
di: Hsu, Tzu-wen, et al.
Pubblicazione: (2025)
di: Hsu, Tzu-wen, et al.
Pubblicazione: (2025)
AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
di: Lin, Tzu-Han, et al.
Pubblicazione: (2025)
di: Lin, Tzu-Han, et al.
Pubblicazione: (2025)
Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025)
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025)
MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
di: Huang, Hsiao-Ying, et al.
Pubblicazione: (2025)
di: Huang, Hsiao-Ying, et al.
Pubblicazione: (2025)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
di: Hsu, Chan-Jan, et al.
Pubblicazione: (2026)
di: Hsu, Chan-Jan, et al.
Pubblicazione: (2026)
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
di: Hsiao, Chi-Yuan, et al.
Pubblicazione: (2026)
di: Hsiao, Chi-Yuan, et al.
Pubblicazione: (2026)
Hearing the Order: Investigating Position Bias in Large Audio-Language Models
di: Lin, Yu-Xiang, et al.
Pubblicazione: (2025)
di: Lin, Yu-Xiang, et al.
Pubblicazione: (2025)
SPAR-K: Scheduled Periodic Alternating Early Exit for Spoken Language Models
di: Huang, Hsiao-Ying, et al.
Pubblicazione: (2026)
di: Huang, Hsiao-Ying, et al.
Pubblicazione: (2026)
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2025)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2025)
Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course
di: Chiang, Cheng-Han, et al.
Pubblicazione: (2024)
di: Chiang, Cheng-Han, et al.
Pubblicazione: (2024)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model
di: Lin, Chun-Hsien, et al.
Pubblicazione: (2024)
di: Lin, Chun-Hsien, et al.
Pubblicazione: (2024)
Transferring Textual Preferences to Vision-Language Understanding through Model Merging
di: Li, Chen-An, et al.
Pubblicazione: (2025)
di: Li, Chen-An, et al.
Pubblicazione: (2025)
On the social bias of speech self-supervised models
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025) -
Property Neurons in Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024) -
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022) -
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025) -
DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)