XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kumar, Shashi, Madikeri, Srikanth, Zuluaga-Gomez, Juan, Villatoro-Tello, Esaú, Thorbecke, Iuliia, Motlicek, Petr, E, Manjunath K, Ganapathiraju, Aravind |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
par: Kumar, Shashi, et autres
Publié: (2024)
par: Kumar, Shashi, et autres
Publié: (2024)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
par: Thorbecke, Iuliia, et autres
Publié: (2024)
par: Thorbecke, Iuliia, et autres
Publié: (2024)
Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
par: Kumar, Shashi, et autres
Publié: (2024)
par: Kumar, Shashi, et autres
Publié: (2024)
Effectiveness of Text, Acoustic, and Lattice-based representations in Spoken Language Understanding tasks
par: Villatoro-Tello, Esaú, et autres
Publié: (2022)
par: Villatoro-Tello, Esaú, et autres
Publié: (2022)
Unifying Global and Near-Context Biasing in a Single Trie Pass
par: Thorbecke, Iuliia, et autres
Publié: (2024)
par: Thorbecke, Iuliia, et autres
Publié: (2024)
Text-only adaptation in LLM-based ASR through text denoising
par: Carofilis, Andrés, et autres
Publié: (2026)
par: Carofilis, Andrés, et autres
Publié: (2026)
Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
par: Burdisso, Sergio, et autres
Publié: (2026)
par: Burdisso, Sergio, et autres
Publié: (2026)
Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR
par: Kumar, Shashi, et autres
Publié: (2026)
par: Kumar, Shashi, et autres
Publié: (2026)
Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering
par: Carofilis, Andres, et autres
Publié: (2025)
par: Carofilis, Andres, et autres
Publié: (2025)
Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
par: Rangappa, Pradeep, et autres
Publié: (2025)
par: Rangappa, Pradeep, et autres
Publié: (2025)
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
par: Kumar, Shashi, et autres
Publié: (2025)
par: Kumar, Shashi, et autres
Publié: (2025)
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
par: Farhadipour, Aref, et autres
Publié: (2025)
par: Farhadipour, Aref, et autres
Publié: (2025)
A Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal Transport
par: Kaloga, Yacouba, et autres
Publié: (2025)
par: Kaloga, Yacouba, et autres
Publié: (2025)
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
par: Baroudi, Séverin, et autres
Publié: (2026)
par: Baroudi, Séverin, et autres
Publié: (2026)
CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
par: Zhao, Wenbo, et autres
Publié: (2024)
par: Zhao, Wenbo, et autres
Publié: (2024)
When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews
par: Watawana, Hasindri, et autres
Publié: (2026)
par: Watawana, Hasindri, et autres
Publié: (2026)
Node-weighted Graph Convolutional Network for Depression Detection in Transcribed Clinical Interviews
par: Burdisso, Sergio, et autres
Publié: (2023)
par: Burdisso, Sergio, et autres
Publié: (2023)
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
par: Farhadipour, Aref, et autres
Publié: (2026)
par: Farhadipour, Aref, et autres
Publié: (2026)
Promptformer: Prompted Conformer Transducer for ASR
par: Duarte-Torres, Sergio, et autres
Publié: (2024)
par: Duarte-Torres, Sergio, et autres
Publié: (2024)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
par: Moriya, Takafumi, et autres
Publié: (2025)
par: Moriya, Takafumi, et autres
Publié: (2025)
Unifying Streaming and Non-streaming Zipformer-based ASR
par: Sharma, Bidisha, et autres
Publié: (2025)
par: Sharma, Bidisha, et autres
Publié: (2025)
XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection
par: Xiao, Yang, et autres
Publié: (2024)
par: Xiao, Yang, et autres
Publié: (2024)
XLSR-MamBo: Scaling the Hybrid Mamba-Attention Backbone for Audio Deepfake Detection
par: Ng, Kwok-Ho, et autres
Publié: (2026)
par: Ng, Kwok-Ho, et autres
Publié: (2026)
Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
par: Farhadipour, Aref, et autres
Publié: (2026)
par: Farhadipour, Aref, et autres
Publié: (2026)
Improving endpoint detection in end-to-end streaming ASR for conversational speech
par: C, Anandh, et autres
Publié: (2025)
par: C, Anandh, et autres
Publié: (2025)
Self-Supervised Learning for Multi-Channel Neural Transducer
par: Kojima, Atsushi
Publié: (2024)
par: Kojima, Atsushi
Publié: (2024)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
par: Andrusenko, Andrei, et autres
Publié: (2026)
par: Andrusenko, Andrei, et autres
Publié: (2026)
Alignment-Free Training for Transducer-based Multi-Talker ASR
par: Moriya, Takafumi, et autres
Publié: (2024)
par: Moriya, Takafumi, et autres
Publié: (2024)
CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
par: Farhadipour, Aref, et autres
Publié: (2025)
par: Farhadipour, Aref, et autres
Publié: (2025)
Learning When to Trust Which Teacher for Weakly Supervised ASR
par: Agrawal, Aakriti, et autres
Publié: (2023)
par: Agrawal, Aakriti, et autres
Publié: (2023)
A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR
par: Biswas, Swadhin, et autres
Publié: (2025)
par: Biswas, Swadhin, et autres
Publié: (2025)
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
par: Tsai, Kai-Chen, et autres
Publié: (2026)
par: Tsai, Kai-Chen, et autres
Publié: (2026)
XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
par: Dat, Phuong Tuan, et autres
Publié: (2025)
par: Dat, Phuong Tuan, et autres
Publié: (2025)
Leveraging ASR Pretrained Conformers for Speaker Verification through Transfer Learning and Knowledge Distillation
par: Cai, Danwei, et autres
Publié: (2023)
par: Cai, Danwei, et autres
Publié: (2023)
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
par: Lee, Hyeonseung, et autres
Publié: (2024)
par: Lee, Hyeonseung, et autres
Publié: (2024)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
par: He, Xiluo, et autres
Publié: (2025)
par: He, Xiluo, et autres
Publié: (2025)
Leveraging Self-Supervised Audio-Visual Pretrained Models to Improve Vocoded Speech Intelligibility in Cochlear Implant Simulation
par: Lai, Richard Lee, et autres
Publié: (2023)
par: Lai, Richard Lee, et autres
Publié: (2023)
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
par: Pražák, Aleš, et autres
Publié: (2025)
par: Pražák, Aleš, et autres
Publié: (2025)
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
par: Han, Jiangyu, et autres
Publié: (2025)
par: Han, Jiangyu, et autres
Publié: (2025)
EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
par: Srivastava, Tejes, et autres
Publié: (2023)
par: Srivastava, Tejes, et autres
Publié: (2023)
Documents similaires
-
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
par: Kumar, Shashi, et autres
Publié: (2024) -
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
par: Thorbecke, Iuliia, et autres
Publié: (2024) -
Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
par: Kumar, Shashi, et autres
Publié: (2024) -
Effectiveness of Text, Acoustic, and Lattice-based representations in Spoken Language Understanding tasks
par: Villatoro-Tello, Esaú, et autres
Publié: (2022) -
Unifying Global and Near-Context Biasing in a Single Trie Pass
par: Thorbecke, Iuliia, et autres
Publié: (2024)