Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuhang, Yang, Yizhou, Peng, Chng, Eng Siong, Zhong, Xionghu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic Speech Separation
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
von: Li, Yuang, et al.
Veröffentlicht: (2024)
von: Li, Yuang, et al.
Veröffentlicht: (2024)
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
von: Liu, Changsong, et al.
Veröffentlicht: (2025)
von: Liu, Changsong, et al.
Veröffentlicht: (2025)
Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
Context-Enhanced Granular Edit Representation for Efficient and Accurate ASR Post-editing
von: Vejsiu, Luan, et al.
Veröffentlicht: (2025)
von: Vejsiu, Luan, et al.
Veröffentlicht: (2025)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Advocating Character Error Rate for Multilingual ASR Evaluation
von: K, Thennal D, et al.
Veröffentlicht: (2024)
von: K, Thennal D, et al.
Veröffentlicht: (2024)
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
von: He, Haorui, et al.
Veröffentlicht: (2024)
von: He, Haorui, et al.
Veröffentlicht: (2024)
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2025)
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2025)
Noise-Aware Speech Separation with Contrastive Learning
von: Zhang, Zizheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zizheng, et al.
Veröffentlicht: (2023)
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission
von: Thakur, Nirmalya Mallick, et al.
Veröffentlicht: (2025)
von: Thakur, Nirmalya Mallick, et al.
Veröffentlicht: (2025)
Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR
von: Mei, Yuxiang, et al.
Veröffentlicht: (2026)
von: Mei, Yuxiang, et al.
Veröffentlicht: (2026)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
von: Rackauckas, Zackary, et al.
Veröffentlicht: (2025)
von: Rackauckas, Zackary, et al.
Veröffentlicht: (2025)
Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
von: Tang, Yun, et al.
Veröffentlicht: (2025)
von: Tang, Yun, et al.
Veröffentlicht: (2025)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
Performant ASR Models for Medical Entities in Accented Speech
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
Crossmodal ASR Error Correction with Discrete Speech Units
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
NTU Speechlab LLM-Based Multilingual ASR System for Interspeech MLC-SLM Challenge 2025
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic Speech Separation
von: Chen, Weiguang, et al.
Veröffentlicht: (2025) -
Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
von: Peng, Yizhou, et al.
Veröffentlicht: (2025) -
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024) -
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
von: Chen, Chen, et al.
Veröffentlicht: (2024) -
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)