Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumar, Shashi, Thorbecke, Iuliia, Burdisso, Sergio, Villatoro-Tello, Esaú, E, Manjunath K, Hacioğlu, Kadri, Rangappa, Pradeep, Motlicek, Petr, Ganapathiraju, Aravind, Stolcke, Andreas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
Text-only adaptation in LLM-based ASR through text denoising
von: Carofilis, Andrés, et al.
Veröffentlicht: (2026)
von: Carofilis, Andrés, et al.
Veröffentlicht: (2026)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
von: Burdisso, Sergio, et al.
Veröffentlicht: (2026)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2026)
Unifying Global and Near-Context Biasing in a Single Trie Pass
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering
von: Carofilis, Andres, et al.
Veröffentlicht: (2025)
von: Carofilis, Andres, et al.
Veröffentlicht: (2025)
Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
von: Rangappa, Pradeep, et al.
Veröffentlicht: (2025)
von: Rangappa, Pradeep, et al.
Veröffentlicht: (2025)
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
von: Kumar, Shashi, et al.
Veröffentlicht: (2025)
von: Kumar, Shashi, et al.
Veröffentlicht: (2025)
Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR
von: Kumar, Shashi, et al.
Veröffentlicht: (2026)
von: Kumar, Shashi, et al.
Veröffentlicht: (2026)
Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
Node-weighted Graph Convolutional Network for Depression Detection in Transcribed Clinical Interviews
von: Burdisso, Sergio, et al.
Veröffentlicht: (2023)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2023)
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
von: Burdisso, Sergio, et al.
Veröffentlicht: (2024)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2024)
Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions
von: Sánchez-Cortés, Dairazalia, et al.
Veröffentlicht: (2024)
von: Sánchez-Cortés, Dairazalia, et al.
Veröffentlicht: (2024)
Effectiveness of Text, Acoustic, and Lattice-based representations in Spoken Language Understanding tasks
von: Villatoro-Tello, Esaú, et al.
Veröffentlicht: (2022)
von: Villatoro-Tello, Esaú, et al.
Veröffentlicht: (2022)
SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
von: Hacioglu, Kadri, et al.
Veröffentlicht: (2025)
von: Hacioglu, Kadri, et al.
Veröffentlicht: (2025)
Slot Filling as a Reasoning Task for SpeechLLMs
von: Hacioglu, Kadri, et al.
Veröffentlicht: (2025)
von: Hacioglu, Kadri, et al.
Veröffentlicht: (2025)
Improving endpoint detection in end-to-end streaming ASR for conversational speech
von: C, Anandh, et al.
Veröffentlicht: (2025)
von: C, Anandh, et al.
Veröffentlicht: (2025)
DAIC-WOZ: On the Validity of Using the Therapist's prompts in Automatic Depression Detection from Clinical Interviews
von: Burdisso, Sergio, et al.
Veröffentlicht: (2024)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2024)
IDIAPers @ Causal News Corpus 2022: Efficient Causal Relation Identification Through a Prompt-based Few-shot Approach
von: Burdisso, Sergio, et al.
Veröffentlicht: (2022)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2022)
A Probabilistic Method for Ranking Refinementin Geographic Information Retrieval
von: Esaú Villatoro-Tello
Veröffentlicht: (2010)
von: Esaú Villatoro-Tello
Veröffentlicht: (2010)
When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews
von: Watawana, Hasindri, et al.
Veröffentlicht: (2026)
von: Watawana, Hasindri, et al.
Veröffentlicht: (2026)
Dialog2Flow: Pre-training Soft-Contrastive Action-Driven Sentence Embeddings for Automatic Dialog Flow Extraction
von: Burdisso, Sergio, et al.
Veröffentlicht: (2024)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2024)
Wealth Accumulation: The Good, the Bad and the Ugly
von: Stewart Lansley
Veröffentlicht: (2024)
von: Stewart Lansley
Veröffentlicht: (2024)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
von: Burdisso, Sergio, et al.
Veröffentlicht: (2025)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2025)
The Cumulants Expansion Approach: The Good, The Bad and The Ugly
von: Kerber, Johannes, et al.
Veröffentlicht: (2025)
von: Kerber, Johannes, et al.
Veröffentlicht: (2025)
Game-Theoretic Cybersecurity: the Good, the Bad and the Ugly
von: Collins, Brandon, et al.
Veröffentlicht: (2024)
von: Collins, Brandon, et al.
Veröffentlicht: (2024)
Coherent Ising Machines: The Good, The Bad, The Ugly
von: Khosravi, Farhad, et al.
Veröffentlicht: (2025)
von: Khosravi, Farhad, et al.
Veröffentlicht: (2025)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
von: Prakash, Jeena, et al.
Veröffentlicht: (2025)
von: Prakash, Jeena, et al.
Veröffentlicht: (2025)
Building EcoCommons Australia - the Good, the Bad and the Ugly
von: Bayraktarov, Elisa
Veröffentlicht: (2023)
von: Bayraktarov, Elisa
Veröffentlicht: (2023)
The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction
von: Wan, Shu, et al.
Veröffentlicht: (2026)
von: Wan, Shu, et al.
Veröffentlicht: (2026)
Vertical Federated Learning in Practice: The Good, the Bad, and the Ugly
von: Wu, Zhaomin, et al.
Veröffentlicht: (2025)
von: Wu, Zhaomin, et al.
Veröffentlicht: (2025)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
von: Burdisso, Sergio, et al.
Veröffentlicht: (2025)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2025)
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
von: Baroudi, Séverin, et al.
Veröffentlicht: (2026)
von: Baroudi, Séverin, et al.
Veröffentlicht: (2026)
List Recoverable Codes: The Good, the Bad, and the Unknown (hopefully not Ugly)
von: Resch, Nicolas, et al.
Veröffentlicht: (2025)
von: Resch, Nicolas, et al.
Veröffentlicht: (2025)
The Good, the Bad, and the Ugly of Atomistic Learning for "Clusters-to-Bulk" Generalization
von: Gawkowski, Mikołaj J., et al.
Veröffentlicht: (2025)
von: Gawkowski, Mikołaj J., et al.
Veröffentlicht: (2025)
Characterizing FaaS Workflows on Public Clouds: The Good, the Bad and the Ugly
von: Kulkarni, Varad, et al.
Veröffentlicht: (2025)
von: Kulkarni, Varad, et al.
Veröffentlicht: (2025)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
von: Woisetschläger, Herbert, et al.
Veröffentlicht: (2023)
von: Woisetschläger, Herbert, et al.
Veröffentlicht: (2023)
Simulation, Modelling and Classification of Wiki Contributors: Spotting The Good, The Bad, and The Ugly
von: Méndez, Silvia García, et al.
Veröffentlicht: (2024)
von: Méndez, Silvia García, et al.
Veröffentlicht: (2024)
Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly
von: Dong, Lianghan, et al.
Veröffentlicht: (2025)
von: Dong, Lianghan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024) -
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
von: Kumar, Shashi, et al.
Veröffentlicht: (2024) -
Text-only adaptation in LLM-based ASR through text denoising
von: Carofilis, Andrés, et al.
Veröffentlicht: (2026) -
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
von: Kumar, Shashi, et al.
Veröffentlicht: (2024) -
Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
von: Burdisso, Sergio, et al.
Veröffentlicht: (2026)