TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Jing, Wang, Chenghao, Yang, Yi, Qian, Lirong, Li, Junjie, Xi, Yu, Wang, Shuai, Yu, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TASU: Text-Only Alignment for Speech Understanding
von: Peng, Jing, et al.
Veröffentlicht: (2025)
von: Peng, Jing, et al.
Veröffentlicht: (2025)
Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning
von: Fang, Yangui, et al.
Veröffentlicht: (2025)
von: Fang, Yangui, et al.
Veröffentlicht: (2025)
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
von: Eom, SooHwan, et al.
Veröffentlicht: (2026)
von: Eom, SooHwan, et al.
Veröffentlicht: (2026)
MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
von: Li, Longhao, et al.
Veröffentlicht: (2025)
von: Li, Longhao, et al.
Veröffentlicht: (2025)
A Survey on Speech Large Language Models for Understanding
von: Peng, Jing, et al.
Veröffentlicht: (2024)
von: Peng, Jing, et al.
Veröffentlicht: (2024)
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
What Does the Speaker Embedding Encode?
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
Joint decoding method for controllable contextual speech recognition based on Speech LLM
von: Fang, Yangui, et al.
Veröffentlicht: (2025)
von: Fang, Yangui, et al.
Veröffentlicht: (2025)
Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
von: Li, Haoyu, et al.
Veröffentlicht: (2025)
von: Li, Haoyu, et al.
Veröffentlicht: (2025)
LV-CTC: Non-autoregressive ASR with CTC and latent variable models
von: Fujita, Yuya, et al.
Veröffentlicht: (2024)
von: Fujita, Yuya, et al.
Veröffentlicht: (2024)
TC-BiMamba: Trans-Chunk bidirectionally within BiMamba for unified streaming and non-streaming ASR
von: She, Qingshun, et al.
Veröffentlicht: (2026)
von: She, Qingshun, et al.
Veröffentlicht: (2026)
Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
von: Xiao, Yunchong, et al.
Veröffentlicht: (2026)
von: Xiao, Yunchong, et al.
Veröffentlicht: (2026)
WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Text-aware Speech Separation for Multi-talker Keyword Spotting
von: Li, Haoyu, et al.
Veröffentlicht: (2024)
von: Li, Haoyu, et al.
Veröffentlicht: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
A Lightweight Hybrid Dual Channel Speech Enhancement System under Low-SNR Conditions
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026)
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026)
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
von: Eom, SooHwan, et al.
Veröffentlicht: (2024)
von: Eom, SooHwan, et al.
Veröffentlicht: (2024)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2024)
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition with BERT and CTC Transformers: A Review
von: Djeffal, Noussaiba, et al.
Veröffentlicht: (2024)
von: Djeffal, Noussaiba, et al.
Veröffentlicht: (2024)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Unimodal Aggregation for CTC-based Speech Recognition
von: Fang, Ying, et al.
Veröffentlicht: (2023)
von: Fang, Ying, et al.
Veröffentlicht: (2023)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation Simulation
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
Hierarchical Control of Emotion Rendering in Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Detect, Attend and Extract: Keyword Guided Target Speaker Extraction
von: Li, Haoyu, et al.
Veröffentlicht: (2026)
von: Li, Haoyu, et al.
Veröffentlicht: (2026)
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
von: Le, Khanh, et al.
Veröffentlicht: (2025)
von: Le, Khanh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TASU: Text-Only Alignment for Speech Understanding
von: Peng, Jing, et al.
Veröffentlicht: (2025) -
Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning
von: Fang, Yangui, et al.
Veröffentlicht: (2025) -
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
von: Eom, SooHwan, et al.
Veröffentlicht: (2026) -
MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
von: Li, Junjie, et al.
Veröffentlicht: (2025) -
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)