WavLLM: Towards Robust and Adaptive Speech Large Language Model
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hu, Shujie, Zhou, Long, Liu, Shujie, Chen, Sanyuan, Meng, Lingwei, Hao, Hongkun, Pan, Jing, Liu, Xunying, Li, Jinyu, Sivasankaran, Sunit, Liu, Linquan, Wei, Furu |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Boosting Large Language Model for Speech Synthesis: An Empirical Study
par: Hao, Hongkun, et autres
Publié: (2023)
par: Hao, Hongkun, et autres
Publié: (2023)
Autoregressive Speech Synthesis without Vector Quantization
par: Meng, Lingwei, et autres
Publié: (2024)
par: Meng, Lingwei, et autres
Publié: (2024)
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
par: Han, Bing, et autres
Publié: (2024)
par: Han, Bing, et autres
Publié: (2024)
WavMark: Watermarking for Audio Generation
par: Chen, Guangyu, et autres
Publié: (2023)
par: Chen, Guangyu, et autres
Publié: (2023)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
par: Chen, Sanyuan, et autres
Publié: (2024)
par: Chen, Sanyuan, et autres
Publié: (2024)
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
par: Niu, Zhikang, et autres
Publié: (2024)
par: Niu, Zhikang, et autres
Publié: (2024)
Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
par: Li, Zhaoqing, et autres
Publié: (2025)
par: Li, Zhaoqing, et autres
Publié: (2025)
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
par: Sivasankaran, Sunit, et autres
Publié: (2024)
par: Sivasankaran, Sunit, et autres
Publié: (2024)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
par: Wang, Hui, et autres
Publié: (2025)
par: Wang, Hui, et autres
Publié: (2025)
Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
par: Meng, Lingwei, et autres
Publié: (2024)
par: Meng, Lingwei, et autres
Publié: (2024)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
par: Wang, Huimeng, et autres
Publié: (2024)
par: Wang, Huimeng, et autres
Publié: (2024)
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
par: Zhong, Tao, et autres
Publié: (2025)
par: Zhong, Tao, et autres
Publié: (2025)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
par: Yuan, Ze, et autres
Publié: (2024)
par: Yuan, Ze, et autres
Publié: (2024)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
par: Chen, Youjun, et autres
Publié: (2025)
par: Chen, Youjun, et autres
Publié: (2025)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
par: Wang, Huimeng, et autres
Publié: (2025)
par: Wang, Huimeng, et autres
Publié: (2025)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
par: HU, Shujie, et autres
Publié: (2025)
par: HU, Shujie, et autres
Publié: (2025)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
par: Gong, Xun, et autres
Publié: (2024)
par: Gong, Xun, et autres
Publié: (2024)
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
par: Wang, Hui, et autres
Publié: (2025)
par: Wang, Hui, et autres
Publié: (2025)
Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition
par: Hu, Shujie, et autres
Publié: (2024)
par: Hu, Shujie, et autres
Publié: (2024)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
par: Li, Guinan, et autres
Publié: (2024)
par: Li, Guinan, et autres
Publié: (2024)
COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
par: Pan, Jing, et autres
Publié: (2023)
par: Pan, Jing, et autres
Publié: (2023)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
par: Kang, Jiawen, et autres
Publié: (2024)
par: Kang, Jiawen, et autres
Publié: (2024)
FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
par: Wang, Hui, et autres
Publié: (2025)
par: Wang, Hui, et autres
Publié: (2025)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
par: Meng, Lingwei, et autres
Publié: (2024)
par: Meng, Lingwei, et autres
Publié: (2024)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
par: Wang, Xiaofei, et autres
Publié: (2023)
par: Wang, Xiaofei, et autres
Publié: (2023)
Closing the Modality Reasoning Gap for Speech Large Language Models
par: Wang, Chaoren, et autres
Publié: (2026)
par: Wang, Chaoren, et autres
Publié: (2026)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
par: Choi, Jeongsoo, et autres
Publié: (2024)
par: Choi, Jeongsoo, et autres
Publié: (2024)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
par: Hu, Shujie, et autres
Publié: (2024)
par: Hu, Shujie, et autres
Publié: (2024)
Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
par: Xie, Xurong, et autres
Publié: (2022)
par: Xie, Xurong, et autres
Publié: (2022)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
par: Xu, Haoning, et autres
Publié: (2025)
par: Xu, Haoning, et autres
Publié: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
par: Li, Jiaqi, et autres
Publié: (2024)
par: Li, Jiaqi, et autres
Publié: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
par: Cui, Mingyu, et autres
Publié: (2024)
par: Cui, Mingyu, et autres
Publié: (2024)
A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation
par: Pei, Hanchen, et autres
Publié: (2026)
par: Pei, Hanchen, et autres
Publié: (2026)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
par: Xie, Xurong, et autres
Publié: (2022)
par: Xie, Xurong, et autres
Publié: (2022)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
par: Kang, Jiawen, et autres
Publié: (2024)
par: Kang, Jiawen, et autres
Publié: (2024)
One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
par: Li, Zhaoqing, et autres
Publié: (2024)
par: Li, Zhaoqing, et autres
Publié: (2024)
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
par: Sun, Zhaokai, et autres
Publié: (2025)
par: Sun, Zhaokai, et autres
Publié: (2025)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
par: Chen, Xueyuan, et autres
Publié: (2024)
par: Chen, Xueyuan, et autres
Publié: (2024)
Towards Effective and Efficient Non-autoregressive Decoding Using Block-based Attention Mask
par: Wang, Tianzi, et autres
Publié: (2024)
par: Wang, Tianzi, et autres
Publié: (2024)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
par: Attia, Ahmed Adel, et autres
Publié: (2024)
par: Attia, Ahmed Adel, et autres
Publié: (2024)
Documents similaires
-
Boosting Large Language Model for Speech Synthesis: An Empirical Study
par: Hao, Hongkun, et autres
Publié: (2023) -
Autoregressive Speech Synthesis without Vector Quantization
par: Meng, Lingwei, et autres
Publié: (2024) -
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
par: Han, Bing, et autres
Publié: (2024) -
WavMark: Watermarking for Audio Generation
par: Chen, Guangyu, et autres
Publié: (2023) -
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
par: Chen, Sanyuan, et autres
Publié: (2024)