TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Shashi, Madikeri, Srikanth, Zuluaga-Gomez, Juan, Thorbecke, Iuliia, Villatoro-Tello, Esaú, Burdisso, Sergio, Motlicek, Petr, Pandia, Karthik, Ganapathiraju, Aravind |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
by: Thorbecke, Iuliia, et al.
Published: (2024)
by: Thorbecke, Iuliia, et al.
Published: (2024)
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
by: Kumar, Shashi, et al.
Published: (2025)
by: Kumar, Shashi, et al.
Published: (2025)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
by: Kumar, Shashi, et al.
Published: (2024)
by: Kumar, Shashi, et al.
Published: (2024)
Node-weighted Graph Convolutional Network for Depression Detection in Transcribed Clinical Interviews
by: Burdisso, Sergio, et al.
Published: (2023)
by: Burdisso, Sergio, et al.
Published: (2023)
Unifying Global and Near-Context Biasing in a Single Trie Pass
by: Thorbecke, Iuliia, et al.
Published: (2024)
by: Thorbecke, Iuliia, et al.
Published: (2024)
Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
by: Kumar, Shashi, et al.
Published: (2024)
by: Kumar, Shashi, et al.
Published: (2024)
Effectiveness of Text, Acoustic, and Lattice-based representations in Spoken Language Understanding tasks
by: Villatoro-Tello, Esaú, et al.
Published: (2022)
by: Villatoro-Tello, Esaú, et al.
Published: (2022)
Dialog2Flow: Pre-training Soft-Contrastive Action-Driven Sentence Embeddings for Automatic Dialog Flow Extraction
by: Burdisso, Sergio, et al.
Published: (2024)
by: Burdisso, Sergio, et al.
Published: (2024)
Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR
by: Kumar, Shashi, et al.
Published: (2026)
by: Kumar, Shashi, et al.
Published: (2026)
Text-only adaptation in LLM-based ASR through text denoising
by: Carofilis, Andrés, et al.
Published: (2026)
by: Carofilis, Andrés, et al.
Published: (2026)
Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
by: Burdisso, Sergio, et al.
Published: (2026)
by: Burdisso, Sergio, et al.
Published: (2026)
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
by: Burdisso, Sergio, et al.
Published: (2024)
by: Burdisso, Sergio, et al.
Published: (2024)
Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions
by: Sánchez-Cortés, Dairazalia, et al.
Published: (2024)
by: Sánchez-Cortés, Dairazalia, et al.
Published: (2024)
Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering
by: Carofilis, Andres, et al.
Published: (2025)
by: Carofilis, Andres, et al.
Published: (2025)
Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
by: Rangappa, Pradeep, et al.
Published: (2025)
by: Rangappa, Pradeep, et al.
Published: (2025)
IDIAPers @ Causal News Corpus 2022: Efficient Causal Relation Identification Through a Prompt-based Few-shot Approach
by: Burdisso, Sergio, et al.
Published: (2022)
by: Burdisso, Sergio, et al.
Published: (2022)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
by: Burdisso, Sergio, et al.
Published: (2025)
by: Burdisso, Sergio, et al.
Published: (2025)
Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
by: Burdisso, Sergio, et al.
Published: (2025)
by: Burdisso, Sergio, et al.
Published: (2025)
DAIC-WOZ: On the Validity of Using the Therapist's prompts in Automatic Depression Detection from Clinical Interviews
by: Burdisso, Sergio, et al.
Published: (2024)
by: Burdisso, Sergio, et al.
Published: (2024)
TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space
by: Garibi, Daniel, et al.
Published: (2025)
by: Garibi, Daniel, et al.
Published: (2025)
A Probabilistic Method for Ranking Refinementin Geographic Information Retrieval
by: Esaú Villatoro-Tello
Published: (2010)
by: Esaú Villatoro-Tello
Published: (2010)
Improving endpoint detection in end-to-end streaming ASR for conversational speech
by: C, Anandh, et al.
Published: (2025)
by: C, Anandh, et al.
Published: (2025)
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews
by: Watawana, Hasindri, et al.
Published: (2026)
by: Watawana, Hasindri, et al.
Published: (2026)
Unifying Streaming and Non-streaming Zipformer-based ASR
by: Sharma, Bidisha, et al.
Published: (2025)
by: Sharma, Bidisha, et al.
Published: (2025)
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
by: Baroudi, Séverin, et al.
Published: (2026)
by: Baroudi, Séverin, et al.
Published: (2026)
Evaluation of Automatic Speech Recognition Using Generative Large Language Models
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
A Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal Transport
by: Kaloga, Yacouba, et al.
Published: (2025)
by: Kaloga, Yacouba, et al.
Published: (2025)
Latent Space Factorization in LoRA
by: Kumar, Shashi, et al.
Published: (2025)
by: Kumar, Shashi, et al.
Published: (2025)
Geometric Latent Reasoning Induces Shorter Generations in LLMs
by: Kumar, Shashi, et al.
Published: (2026)
by: Kumar, Shashi, et al.
Published: (2026)
Temporal fine-tuning for early risk detection
by: Thompson, Horacio, et al.
Published: (2025)
by: Thompson, Horacio, et al.
Published: (2025)
When De-noising Hurts: A Systematic Study of Speech Enhancement Effects on Modern Medical ASR Systems
by: Chondhekar, Sujal, et al.
Published: (2025)
by: Chondhekar, Sujal, et al.
Published: (2025)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
by: Andrusenko, Andrei, et al.
Published: (2026)
by: Andrusenko, Andrei, et al.
Published: (2026)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
by: Moriya, Takafumi, et al.
Published: (2025)
by: Moriya, Takafumi, et al.
Published: (2025)
Promptformer: Prompted Conformer Transducer for ASR
by: Duarte-Torres, Sergio, et al.
Published: (2024)
by: Duarte-Torres, Sergio, et al.
Published: (2024)
Typhoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition
by: Sirichotedumrong, Warit, et al.
Published: (2026)
by: Sirichotedumrong, Warit, et al.
Published: (2026)
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
by: Farhadipour, Aref, et al.
Published: (2026)
by: Farhadipour, Aref, et al.
Published: (2026)
Measurement of Social Well-being and Progress
by: Thorbecke, Erik
by: Thorbecke, Erik
El probleme del empleo: evaluación crítica de informes de la OIT sobre cuatro países
by: Erik Thorbecke
Published: (1973)
by: Erik Thorbecke
Published: (1973)
Similar Items
-
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
by: Thorbecke, Iuliia, et al.
Published: (2024) -
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
by: Kumar, Shashi, et al.
Published: (2025) -
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
by: Kumar, Shashi, et al.
Published: (2024) -
Node-weighted Graph Convolutional Network for Depression Detection in Transcribed Clinical Interviews
by: Burdisso, Sergio, et al.
Published: (2023) -
Unifying Global and Near-Context Biasing in a Single Trie Pass
by: Thorbecke, Iuliia, et al.
Published: (2024)