CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Xueyuan, Yang, Dongchao, Wang, Dingdong, Wu, Xixin, Wu, Zhiyong, Meng, Helen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
di: Chen, Xueyuan, et al.
Pubblicazione: (2025)
di: Chen, Xueyuan, et al.
Pubblicazione: (2025)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
di: Chen, Xueyuan, et al.
Pubblicazione: (2024)
di: Chen, Xueyuan, et al.
Pubblicazione: (2024)
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
di: Wang, Yuejiao, et al.
Pubblicazione: (2024)
di: Wang, Yuejiao, et al.
Pubblicazione: (2024)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
di: Guo, Haohan, et al.
Pubblicazione: (2024)
di: Guo, Haohan, et al.
Pubblicazione: (2024)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
di: Guo, Haohan, et al.
Pubblicazione: (2024)
di: Guo, Haohan, et al.
Pubblicazione: (2024)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
di: Wang, Dingdong, et al.
Pubblicazione: (2024)
di: Wang, Dingdong, et al.
Pubblicazione: (2024)
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
di: Wang, Yuanyuan, et al.
Pubblicazione: (2026)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2026)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
di: Wu, Wenxuan, et al.
Pubblicazione: (2024)
di: Wu, Wenxuan, et al.
Pubblicazione: (2024)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
di: Wang, Yuanyuan, et al.
Pubblicazione: (2024)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2024)
Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
di: Guo, Haohan, et al.
Pubblicazione: (2024)
di: Guo, Haohan, et al.
Pubblicazione: (2024)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
UniSep: Universal Target Audio Separation with Language Models at Scale
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
di: Lei, Shun, et al.
Pubblicazione: (2023)
di: Lei, Shun, et al.
Pubblicazione: (2023)
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
Bridging the gap between training and inference in LM-based TTS models
di: Zhang, Ruonan, et al.
Pubblicazione: (2025)
di: Zhang, Ruonan, et al.
Pubblicazione: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
di: Xie, Jingran, et al.
Pubblicazione: (2025)
di: Xie, Jingran, et al.
Pubblicazione: (2025)
Personalized Neural Speech Codec
di: Jang, Inseon, et al.
Pubblicazione: (2024)
di: Jang, Inseon, et al.
Pubblicazione: (2024)
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
di: Zhou, Yixuan, et al.
Pubblicazione: (2024)
di: Zhou, Yixuan, et al.
Pubblicazione: (2024)
Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition
di: Hu, Shujie, et al.
Pubblicazione: (2024)
di: Hu, Shujie, et al.
Pubblicazione: (2024)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
di: Li, Jiaqi, et al.
Pubblicazione: (2025)
di: Li, Jiaqi, et al.
Pubblicazione: (2025)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
di: Xin, Detai, et al.
Pubblicazione: (2024)
di: Xin, Detai, et al.
Pubblicazione: (2024)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
di: HU, Shujie, et al.
Pubblicazione: (2025)
di: HU, Shujie, et al.
Pubblicazione: (2025)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
di: Chen, Weidong, et al.
Pubblicazione: (2025)
di: Chen, Weidong, et al.
Pubblicazione: (2025)
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024)
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024)
SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
MuCodec: Ultra Low-Bitrate Music Codec
di: Xu, Yaoxun, et al.
Pubblicazione: (2024)
di: Xu, Yaoxun, et al.
Pubblicazione: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
Probing Whisper for Dysarthric Speech in Detection and Assessment
di: Yue, Zhengjun, et al.
Pubblicazione: (2025)
di: Yue, Zhengjun, et al.
Pubblicazione: (2025)
Probing the Robustness Properties of Neural Speech Codecs
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)
SpatialCodec: Neural Spatial Speech Coding
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2023)
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2023)
A Neural Speech Codec for Noise Robust Speech Coding
di: Huang, Jiayi, et al.
Pubblicazione: (2023)
di: Huang, Jiayi, et al.
Pubblicazione: (2023)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
di: Chen, Xueyuan, et al.
Pubblicazione: (2025) -
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
di: Chen, Xueyuan, et al.
Pubblicazione: (2024) -
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
di: Wang, Yuejiao, et al.
Pubblicazione: (2024) -
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
di: Yang, Dongchao, et al.
Pubblicazione: (2024) -
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
di: Guo, Haohan, et al.
Pubblicazione: (2024)