Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Mu, Chen, Szu-Jui, Xie, Jiamin, Hansen, John |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
por: Yang, Haoyuan, et al.
Publicado: (2026)
por: Yang, Haoyuan, et al.
Publicado: (2026)
MixRep: Hidden Representation Mixup for Low-Resource Speech Recognition
por: Xie, Jiamin, et al.
Publicado: (2023)
por: Xie, Jiamin, et al.
Publicado: (2023)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
por: Chen, Szu-Jui, et al.
Publicado: (2026)
por: Chen, Szu-Jui, et al.
Publicado: (2026)
DEFORMER: Coupling Deformed Localized Patterns with Global Context for Robust End-to-end Speech Recognition
por: Xie, Jiamin, et al.
Publicado: (2022)
por: Xie, Jiamin, et al.
Publicado: (2022)
Acoustic modeling for Overlapping Speech Recognition: JHU Chime-5 Challenge System
por: Manohar, Vimal, et al.
Publicado: (2024)
por: Manohar, Vimal, et al.
Publicado: (2024)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
por: Chen, Peikun, et al.
Publicado: (2024)
por: Chen, Peikun, et al.
Publicado: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
por: Wang, Shih-heng, et al.
Publicado: (2024)
por: Wang, Shih-heng, et al.
Publicado: (2024)
Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech
por: Yang, Mu, et al.
Publicado: (2026)
por: Yang, Mu, et al.
Publicado: (2026)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
por: Kim, Jongsuk, et al.
Publicado: (2025)
por: Kim, Jongsuk, et al.
Publicado: (2025)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
por: Mu, Bingshen, et al.
Publicado: (2025)
por: Mu, Bingshen, et al.
Publicado: (2025)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
por: Lin, Meng-Ping, et al.
Publicado: (2025)
por: Lin, Meng-Ping, et al.
Publicado: (2025)
Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization
por: Li, Xueqing, et al.
Publicado: (2025)
por: Li, Xueqing, et al.
Publicado: (2025)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
por: Tian, Wenjie, et al.
Publicado: (2026)
por: Tian, Wenjie, et al.
Publicado: (2026)
Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition
por: Ok, Seaone, et al.
Publicado: (2026)
por: Ok, Seaone, et al.
Publicado: (2026)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
por: Wang, Kuan-Chen, et al.
Publicado: (2024)
por: Wang, Kuan-Chen, et al.
Publicado: (2024)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
por: Tian, Jingguang, et al.
Publicado: (2024)
por: Tian, Jingguang, et al.
Publicado: (2024)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
por: Su, Fei, et al.
Publicado: (2026)
por: Su, Fei, et al.
Publicado: (2026)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
por: Ji, Shengpeng, et al.
Publicado: (2024)
por: Ji, Shengpeng, et al.
Publicado: (2024)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
por: Wei, Linye, et al.
Publicado: (2025)
por: Wei, Linye, et al.
Publicado: (2025)
Fairness of Automatic Speech Recognition in Cleft Lip and Palate Speech
por: Bhattacharjee, Susmita, et al.
Publicado: (2025)
por: Bhattacharjee, Susmita, et al.
Publicado: (2025)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
por: Dhawan, Kunal, et al.
Publicado: (2024)
por: Dhawan, Kunal, et al.
Publicado: (2024)
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults
por: Attia, Ahmed Adel, et al.
Publicado: (2023)
por: Attia, Ahmed Adel, et al.
Publicado: (2023)
Discrete Diffusion for Generative Modeling of Text-Aligned Speech Tokens
por: Ku, Pin-Jui, et al.
Publicado: (2025)
por: Ku, Pin-Jui, et al.
Publicado: (2025)
Discrete Audio Representations for Automated Audio Captioning
por: Tian, Jingguang, et al.
Publicado: (2025)
por: Tian, Jingguang, et al.
Publicado: (2025)
USAD: Universal Speech and Audio Representation via Distillation
por: Chang, Heng-Jui, et al.
Publicado: (2025)
por: Chang, Heng-Jui, et al.
Publicado: (2025)
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
por: Aronowitz, Hagai, et al.
Publicado: (2026)
por: Aronowitz, Hagai, et al.
Publicado: (2026)
Interpreting the Role of Visemes in Audio-Visual Speech Recognition
por: Papadopoulos, Aristeidis, et al.
Publicado: (2025)
por: Papadopoulos, Aristeidis, et al.
Publicado: (2025)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
por: Wang, Huimeng, et al.
Publicado: (2025)
por: Wang, Huimeng, et al.
Publicado: (2025)
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
por: Yang, Xiaoyu, et al.
Publicado: (2025)
por: Yang, Xiaoyu, et al.
Publicado: (2025)
Unsupervised Online Continual Learning for Automatic Speech Recognition
por: Eeckt, Steven Vander, et al.
Publicado: (2024)
por: Eeckt, Steven Vander, et al.
Publicado: (2024)
Using Songs to Improve Kazakh Automatic Speech Recognition
por: Yeshpanov, Rustem
Publicado: (2026)
por: Yeshpanov, Rustem
Publicado: (2026)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
por: Ku, Pin-Jui, et al.
Publicado: (2024)
por: Ku, Pin-Jui, et al.
Publicado: (2024)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
por: Xue, Hongfei, et al.
Publicado: (2023)
por: Xue, Hongfei, et al.
Publicado: (2023)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
por: Bai, Ye, et al.
Publicado: (2024)
por: Bai, Ye, et al.
Publicado: (2024)
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
por: Li, Yangze, et al.
Publicado: (2024)
por: Li, Yangze, et al.
Publicado: (2024)
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
por: Cui, Zhongjian, et al.
Publicado: (2025)
por: Cui, Zhongjian, et al.
Publicado: (2025)
Multi-Distillation from Speech and Music Representation Models
por: Wei, Jui-Chiang, et al.
Publicado: (2025)
por: Wei, Jui-Chiang, et al.
Publicado: (2025)
Non-Intrusive Automatic Speech Recognition Refinement: A Survey
por: Peyghan, Mohammad Reza, et al.
Publicado: (2025)
por: Peyghan, Mohammad Reza, et al.
Publicado: (2025)
Improving Automatic Speech Recognition for Speakers Treated for Oral Cancer using Data Augmentation and LLM Error Correction
por: Folkertsma, Hidde, et al.
Publicado: (2026)
por: Folkertsma, Hidde, et al.
Publicado: (2026)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
por: Wei, Kun, et al.
Publicado: (2023)
por: Wei, Kun, et al.
Publicado: (2023)
Ejemplares similares
-
Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
por: Yang, Haoyuan, et al.
Publicado: (2026) -
MixRep: Hidden Representation Mixup for Low-Resource Speech Recognition
por: Xie, Jiamin, et al.
Publicado: (2023) -
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
por: Chen, Szu-Jui, et al.
Publicado: (2026) -
DEFORMER: Coupling Deformed Localized Patterns with Global Context for Robust End-to-end Speech Recognition
por: Xie, Jiamin, et al.
Publicado: (2022) -
Acoustic modeling for Overlapping Speech Recognition: JHU Chime-5 Challenge System
por: Manohar, Vimal, et al.
Publicado: (2024)