SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Le, Khanh, Ho, Tuan Vu, Tran, Dung, Chau, Duc Thanh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
por: Le, Khanh, et al.
Publicado: (2025)
por: Le, Khanh, et al.
Publicado: (2025)
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
por: Le, Khanh, et al.
Publicado: (2025)
por: Le, Khanh, et al.
Publicado: (2025)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
por: Tsunoo, Emiru, et al.
Publicado: (2023)
por: Tsunoo, Emiru, et al.
Publicado: (2023)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
por: Hou, Junfeng, et al.
Publicado: (2024)
por: Hou, Junfeng, et al.
Publicado: (2024)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
por: Pham, The Hieu, et al.
Publicado: (2025)
por: Pham, The Hieu, et al.
Publicado: (2025)
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization
por: Ho, Luong, et al.
Publicado: (2025)
por: Ho, Luong, et al.
Publicado: (2025)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
por: Bataev, Vladimir
Publicado: (2025)
por: Bataev, Vladimir
Publicado: (2025)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
por: Gong, Xun, et al.
Publicado: (2024)
por: Gong, Xun, et al.
Publicado: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
por: Shen, Siyuan, et al.
Publicado: (2024)
por: Shen, Siyuan, et al.
Publicado: (2024)
Unimodal Aggregation for CTC-based Speech Recognition
por: Fang, Ying, et al.
Publicado: (2023)
por: Fang, Ying, et al.
Publicado: (2023)
Multi-blank Transducers for Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2022)
por: Xu, Hainan, et al.
Publicado: (2022)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
por: Sakuma, Asahi, et al.
Publicado: (2025)
por: Sakuma, Asahi, et al.
Publicado: (2025)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
por: Burchi, Maxime, et al.
Publicado: (2024)
por: Burchi, Maxime, et al.
Publicado: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
por: Du, Chenpeng, et al.
Publicado: (2024)
por: Du, Chenpeng, et al.
Publicado: (2024)
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
por: Yamashita, Natsuo, et al.
Publicado: (2026)
por: Yamashita, Natsuo, et al.
Publicado: (2026)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
por: Dang, Trung, et al.
Publicado: (2024)
por: Dang, Trung, et al.
Publicado: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
por: Kang, Jiawen, et al.
Publicado: (2024)
por: Kang, Jiawen, et al.
Publicado: (2024)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
por: Zhao, Junchuan, et al.
Publicado: (2026)
por: Zhao, Junchuan, et al.
Publicado: (2026)
Stream-based Active Learning for Anomalous Sound Detection in Machine Condition Monitoring
por: Ho, Tuan Vu, et al.
Publicado: (2024)
por: Ho, Tuan Vu, et al.
Publicado: (2024)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
Enhancing CTC-Based Visual Speech Recognition
por: Laux, Hendrik, et al.
Publicado: (2024)
por: Laux, Hendrik, et al.
Publicado: (2024)
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
por: Xi, Yu, et al.
Publicado: (2024)
por: Xi, Yu, et al.
Publicado: (2024)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
por: Zhou, Jiaming, et al.
Publicado: (2023)
por: Zhou, Jiaming, et al.
Publicado: (2023)
Zero-Shot Text-to-Speech from Continuous Text Streams
por: Dang, Trung, et al.
Publicado: (2024)
por: Dang, Trung, et al.
Publicado: (2024)
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
por: Phuong, Tuan Dat, et al.
Publicado: (2025)
por: Phuong, Tuan Dat, et al.
Publicado: (2025)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
por: Nasretdinov, Rauf, et al.
Publicado: (2025)
por: Nasretdinov, Rauf, et al.
Publicado: (2025)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
por: Do, Cong-Thanh, et al.
Publicado: (2024)
por: Do, Cong-Thanh, et al.
Publicado: (2024)
CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
por: Zhang, Tian-Hao, et al.
Publicado: (2023)
por: Zhang, Tian-Hao, et al.
Publicado: (2023)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
por: Nespoli, Francesco, et al.
Publicado: (2024)
por: Nespoli, Francesco, et al.
Publicado: (2024)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
por: Leung, Wing-Zin, et al.
Publicado: (2024)
por: Leung, Wing-Zin, et al.
Publicado: (2024)
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
por: Bondaruk, Łukasz, et al.
Publicado: (2024)
por: Bondaruk, Łukasz, et al.
Publicado: (2024)
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
por: Pusateri, Ernest, et al.
Publicado: (2024)
por: Pusateri, Ernest, et al.
Publicado: (2024)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
por: Andrusenko, Andrei, et al.
Publicado: (2024)
por: Andrusenko, Andrei, et al.
Publicado: (2024)
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
por: Derington, Anna, et al.
Publicado: (2023)
por: Derington, Anna, et al.
Publicado: (2023)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2024)
por: Xu, Hainan, et al.
Publicado: (2024)
Xi+: Uncertainty Supervision for Robust Speaker Embedding
por: Li, Junjie, et al.
Publicado: (2025)
por: Li, Junjie, et al.
Publicado: (2025)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
por: Wang, Shih-heng, et al.
Publicado: (2024)
por: Wang, Shih-heng, et al.
Publicado: (2024)
USM RNN-T model weights binarization
por: Rybakov, Oleg, et al.
Publicado: (2024)
por: Rybakov, Oleg, et al.
Publicado: (2024)
Ejemplares similares
-
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
por: Le, Khanh, et al.
Publicado: (2025) -
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
por: Le, Khanh, et al.
Publicado: (2025) -
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
por: Tsunoo, Emiru, et al.
Publicado: (2023) -
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
por: Hou, Junfeng, et al.
Publicado: (2024) -
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
por: Pham, The Hieu, et al.
Publicado: (2025)