Samba-ASR: State-Of-The-Art Speech Recognition Leveraging Structured State-Space Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shakhadri, Syed Abdul Gaffar, KR, Kruthika, Angadi, Kartik Basavaraj |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
von: Shakhadri, Syed Abdul Gaffar, et al.
Veröffentlicht: (2025)
von: Shakhadri, Syed Abdul Gaffar, et al.
Veröffentlicht: (2025)
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems
von: Ghimire, Rupak Raj, et al.
Veröffentlicht: (2024)
von: Ghimire, Rupak Raj, et al.
Veröffentlicht: (2024)
FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
von: Xu, Kaituo, et al.
Veröffentlicht: (2026)
von: Xu, Kaituo, et al.
Veröffentlicht: (2026)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
von: Srivastav, Vaibhav, et al.
Veröffentlicht: (2025)
von: Srivastav, Vaibhav, et al.
Veröffentlicht: (2025)
ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech Recognition Systems
von: Rai, Anand, et al.
Veröffentlicht: (2025)
von: Rai, Anand, et al.
Veröffentlicht: (2025)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
von: Wang, He, et al.
Veröffentlicht: (2025)
von: Wang, He, et al.
Veröffentlicht: (2025)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Fine-Tuning Small Language Models for Domain-Specific AI: An Edge AI Perspective
von: Aralimatti, Rakshit, et al.
Veröffentlicht: (2025)
von: Aralimatti, Rakshit, et al.
Veröffentlicht: (2025)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks
von: Chen, Sizhou, et al.
Veröffentlicht: (2023)
von: Chen, Sizhou, et al.
Veröffentlicht: (2023)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
Advancing African-Accented Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models
von: Dossou, Bonaventure F. P.
Veröffentlicht: (2023)
von: Dossou, Bonaventure F. P.
Veröffentlicht: (2023)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
ASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages
von: Kumar, Subham, et al.
Veröffentlicht: (2025)
von: Kumar, Subham, et al.
Veröffentlicht: (2025)
A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023)
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023)
Performant ASR Models for Medical Entities in Accented Speech
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
Speech Recognition on TV Series with Video-guided Post-ASR Correction
von: Yang, Haoyuan, et al.
Veröffentlicht: (2025)
von: Yang, Haoyuan, et al.
Veröffentlicht: (2025)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2024)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2024)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
von: Aboeitta, Ahmed, et al.
Veröffentlicht: (2025)
von: Aboeitta, Ahmed, et al.
Veröffentlicht: (2025)
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
von: Xie, Zhifei, et al.
Veröffentlicht: (2026)
von: Xie, Zhifei, et al.
Veröffentlicht: (2026)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
Crossmodal ASR Error Correction with Discrete Speech Units
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
von: Shakhadri, Syed Abdul Gaffar, et al.
Veröffentlicht: (2025) -
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems
von: Ghimire, Rupak Raj, et al.
Veröffentlicht: (2024) -
FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
von: Xu, Kaituo, et al.
Veröffentlicht: (2026) -
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024) -
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)