CUSIDE-array: A Streaming Multi-Channel End-to-End Speech Recognition System with Realistic Evaluations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kong, Xiangzhu, Ning, Tianqi, Huang, Hao, Ou, Zhijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2025)
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2025)
CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024)
Reference Channel Selection by Multi-Channel Masking for End-to-End Multi-Channel Speech Enhancement
von: Dai, Wang, et al.
Veröffentlicht: (2024)
von: Dai, Wang, et al.
Veröffentlicht: (2024)
End-to-End Speech Recognition with Pre-trained Masked Language Model
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2022)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2022)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation
von: Wang, Pu, et al.
Veröffentlicht: (2024)
von: Wang, Pu, et al.
Veröffentlicht: (2024)
Enhancing Fully Formatted End-to-End Speech Recognition with Knowledge Distillation via Multi-Codebook Vector Quantization
von: You, Jian, et al.
Veröffentlicht: (2025)
von: You, Jian, et al.
Veröffentlicht: (2025)
SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription
von: Dai, Yuhang, et al.
Veröffentlicht: (2026)
von: Dai, Yuhang, et al.
Veröffentlicht: (2026)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2024)
von: Chen, Jinming, et al.
Veröffentlicht: (2024)
Breaking Walls: Pioneering Automatic Speech Recognition for Central Kurdish: End-to-End Transformer Paradigm
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios
von: Jing, Kangqi, et al.
Veröffentlicht: (2025)
von: Jing, Kangqi, et al.
Veröffentlicht: (2025)
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
von: Zhuang, Zhuoran, et al.
Veröffentlicht: (2026)
von: Zhuang, Zhuoran, et al.
Veröffentlicht: (2026)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational Autoencoder
von: Cui, Jianwei, et al.
Veröffentlicht: (2024)
von: Cui, Jianwei, et al.
Veröffentlicht: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding
von: Su, Tianhui, et al.
Veröffentlicht: (2026)
von: Su, Tianhui, et al.
Veröffentlicht: (2026)
Data Augmentation for End-to-end Code-switching Speech Recognition
von: Du, Chenpeng, et al.
Veröffentlicht: (2020)
von: Du, Chenpeng, et al.
Veröffentlicht: (2020)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
Chunkwise Aligners for Streaming Speech Recognition
von: Teo, Wen Shen, et al.
Veröffentlicht: (2026)
von: Teo, Wen Shen, et al.
Veröffentlicht: (2026)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
Continual Learning for Monolingual End-to-End Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2021)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2021)
DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
von: Ning, Ziqian, et al.
Veröffentlicht: (2025)
von: Ning, Ziqian, et al.
Veröffentlicht: (2025)
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2026)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2026)
Privacy-Preserving End-to-End Full-Duplex Speech Dialogue Models
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
von: Lin, Wan, et al.
Veröffentlicht: (2024)
von: Lin, Wan, et al.
Veröffentlicht: (2024)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2025) -
CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024) -
Reference Channel Selection by Multi-Channel Masking for End-to-End Multi-Channel Speech Enhancement
von: Dai, Wang, et al.
Veröffentlicht: (2024) -
End-to-End Speech Recognition with Pre-trained Masked Language Model
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024) -
Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2022)