Text-only adaptation in LLM-based ASR through text denoising
Fuente:
arXiv
Saved in:
| Main Authors: | Carofilis, Andrés, Burdisso, Sergio, Villatoro-Tello, Esaú, Kumar, Shashi, Hacioglu, Kadri, Madikeri, Srikanth, Rangappa, Pradeep, E, Manjunath K, Motlicek, Petr, Venkatesan, Shankar, Stolcke, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
by: Kumar, Shashi, et al.
Published: (2025)
by: Kumar, Shashi, et al.
Published: (2025)
Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
by: Burdisso, Sergio, et al.
Published: (2026)
by: Burdisso, Sergio, et al.
Published: (2026)
Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering
by: Carofilis, Andres, et al.
Published: (2025)
by: Carofilis, Andres, et al.
Published: (2025)
Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
by: Rangappa, Pradeep, et al.
Published: (2025)
by: Rangappa, Pradeep, et al.
Published: (2025)
Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
by: Kumar, Shashi, et al.
Published: (2024)
by: Kumar, Shashi, et al.
Published: (2024)
Unifying Global and Near-Context Biasing in a Single Trie Pass
by: Thorbecke, Iuliia, et al.
Published: (2024)
by: Thorbecke, Iuliia, et al.
Published: (2024)
Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR
by: Kumar, Shashi, et al.
Published: (2026)
by: Kumar, Shashi, et al.
Published: (2026)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
by: Kumar, Shashi, et al.
Published: (2024)
by: Kumar, Shashi, et al.
Published: (2024)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
by: Kumar, Shashi, et al.
Published: (2024)
by: Kumar, Shashi, et al.
Published: (2024)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
by: Thorbecke, Iuliia, et al.
Published: (2024)
by: Thorbecke, Iuliia, et al.
Published: (2024)
Node-weighted Graph Convolutional Network for Depression Detection in Transcribed Clinical Interviews
by: Burdisso, Sergio, et al.
Published: (2023)
by: Burdisso, Sergio, et al.
Published: (2023)
Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
by: Prakash, Jeena, et al.
Published: (2025)
by: Prakash, Jeena, et al.
Published: (2025)
Effectiveness of Text, Acoustic, and Lattice-based representations in Spoken Language Understanding tasks
by: Villatoro-Tello, Esaú, et al.
Published: (2022)
by: Villatoro-Tello, Esaú, et al.
Published: (2022)
Dialog2Flow: Pre-training Soft-Contrastive Action-Driven Sentence Embeddings for Automatic Dialog Flow Extraction
by: Burdisso, Sergio, et al.
Published: (2024)
by: Burdisso, Sergio, et al.
Published: (2024)
Improving endpoint detection in end-to-end streaming ASR for conversational speech
by: C, Anandh, et al.
Published: (2025)
by: C, Anandh, et al.
Published: (2025)
When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews
by: Watawana, Hasindri, et al.
Published: (2026)
by: Watawana, Hasindri, et al.
Published: (2026)
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
by: Burdisso, Sergio, et al.
Published: (2024)
by: Burdisso, Sergio, et al.
Published: (2024)
Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions
by: Sánchez-Cortés, Dairazalia, et al.
Published: (2024)
by: Sánchez-Cortés, Dairazalia, et al.
Published: (2024)
Unifying Streaming and Non-streaming Zipformer-based ASR
by: Sharma, Bidisha, et al.
Published: (2025)
by: Sharma, Bidisha, et al.
Published: (2025)
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
by: Baroudi, Séverin, et al.
Published: (2026)
by: Baroudi, Séverin, et al.
Published: (2026)
A Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal Transport
by: Kaloga, Yacouba, et al.
Published: (2025)
by: Kaloga, Yacouba, et al.
Published: (2025)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
by: Bataev, Vladimir, et al.
Published: (2023)
by: Bataev, Vladimir, et al.
Published: (2023)
SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
by: Hacioglu, Kadri, et al.
Published: (2025)
by: Hacioglu, Kadri, et al.
Published: (2025)
Slot Filling as a Reasoning Task for SpeechLLMs
by: Hacioglu, Kadri, et al.
Published: (2025)
by: Hacioglu, Kadri, et al.
Published: (2025)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
by: Burdisso, Sergio, et al.
Published: (2025)
by: Burdisso, Sergio, et al.
Published: (2025)
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
by: Farhadipour, Aref, et al.
Published: (2026)
by: Farhadipour, Aref, et al.
Published: (2026)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
by: Burdisso, Sergio, et al.
Published: (2025)
by: Burdisso, Sergio, et al.
Published: (2025)
DAIC-WOZ: On the Validity of Using the Therapist's prompts in Automatic Depression Detection from Clinical Interviews
by: Burdisso, Sergio, et al.
Published: (2024)
by: Burdisso, Sergio, et al.
Published: (2024)
IDIAPers @ Causal News Corpus 2022: Efficient Causal Relation Identification Through a Prompt-based Few-shot Approach
by: Burdisso, Sergio, et al.
Published: (2022)
by: Burdisso, Sergio, et al.
Published: (2022)
Learning When to Trust Which Teacher for Weakly Supervised ASR
by: Agrawal, Aakriti, et al.
Published: (2023)
by: Agrawal, Aakriti, et al.
Published: (2023)
Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
by: Farhadipour, Aref, et al.
Published: (2026)
by: Farhadipour, Aref, et al.
Published: (2026)
Text adaptation for speaker verification with speaker-text factorized embeddings
by: Yang, Yexin, et al.
Published: (2025)
by: Yang, Yexin, et al.
Published: (2025)
A Probabilistic Method for Ranking Refinementin Geographic Information Retrieval
by: Esaú Villatoro-Tello
Published: (2010)
by: Esaú Villatoro-Tello
Published: (2010)
CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Improving fairness in speaker verification via Group-adapted Fusion Network
by: Shen, Hua, et al.
Published: (2022)
by: Shen, Hua, et al.
Published: (2022)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
by: Chen, Qian, et al.
Published: (2023)
by: Chen, Qian, et al.
Published: (2023)
SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR
by: Shankar, Natarajan Balaji, et al.
Published: (2024)
by: Shankar, Natarajan Balaji, et al.
Published: (2024)
CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR
by: Shankar, Natarajan Balaji, et al.
Published: (2025)
by: Shankar, Natarajan Balaji, et al.
Published: (2025)
Similar Items
-
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
by: Kumar, Shashi, et al.
Published: (2025) -
Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
by: Burdisso, Sergio, et al.
Published: (2026) -
Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering
by: Carofilis, Andres, et al.
Published: (2025) -
Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
by: Rangappa, Pradeep, et al.
Published: (2025) -
Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
by: Kumar, Shashi, et al.
Published: (2024)