In-Materia Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Zolfagharinejad, Mohamadreza, Büchel, Julian, Cassola, Lorenzo, Kinge, Sachin, Syed, Ghazi Sarwat, Sebastian, Abu, van der Wiel, Wilfred G. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
di: Nasretdinov, Rauf, et al.
Pubblicazione: (2025)
di: Nasretdinov, Rauf, et al.
Pubblicazione: (2025)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
di: Yen, Hao, et al.
Pubblicazione: (2024)
di: Yen, Hao, et al.
Pubblicazione: (2024)
Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
di: Talpur, Unzela, et al.
Pubblicazione: (2025)
di: Talpur, Unzela, et al.
Pubblicazione: (2025)
Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
di: Ueda, Lucas H., et al.
Pubblicazione: (2026)
di: Ueda, Lucas H., et al.
Pubblicazione: (2026)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2024)
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2024)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
di: Bai, Ye, et al.
Pubblicazione: (2024)
di: Bai, Ye, et al.
Pubblicazione: (2024)
Speech Recognition for Analysis of Police Radio Communication
di: Srivastava, Tejes, et al.
Pubblicazione: (2024)
di: Srivastava, Tejes, et al.
Pubblicazione: (2024)
Two-pass Endpoint Detection for Speech Recognition
di: Raju, Anirudh, et al.
Pubblicazione: (2024)
di: Raju, Anirudh, et al.
Pubblicazione: (2024)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
Efficient Long Speech Sequence Modelling for Time-Domain Depression Level Estimation
di: Li, Shuanglin, et al.
Pubblicazione: (2025)
di: Li, Shuanglin, et al.
Pubblicazione: (2025)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
di: Lin, Zhaofeng, et al.
Pubblicazione: (2024)
di: Lin, Zhaofeng, et al.
Pubblicazione: (2024)
Attention-Guided Adaptation for Code-Switching Speech Recognition
di: Aditya, Bobbi, et al.
Pubblicazione: (2023)
di: Aditya, Bobbi, et al.
Pubblicazione: (2023)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
Dataset-Distillation Generative Model for Speech Emotion Recognition
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2024)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2024)
THAI Speech Emotion Recognition (THAI-SER) corpus
di: Wongpithayadisai, Jilamika, et al.
Pubblicazione: (2025)
di: Wongpithayadisai, Jilamika, et al.
Pubblicazione: (2025)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
di: Sun, Haoqin, et al.
Pubblicazione: (2024)
di: Sun, Haoqin, et al.
Pubblicazione: (2024)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
di: Chen, Peikun, et al.
Pubblicazione: (2024)
di: Chen, Peikun, et al.
Pubblicazione: (2024)
USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models
di: Ding, Shaojin, et al.
Pubblicazione: (2023)
di: Ding, Shaojin, et al.
Pubblicazione: (2023)
Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language
di: Abu, Turi, et al.
Pubblicazione: (2025)
di: Abu, Turi, et al.
Pubblicazione: (2025)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
di: Wagner, Dominik, et al.
Pubblicazione: (2025)
di: Wagner, Dominik, et al.
Pubblicazione: (2025)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
di: Zheng, Xiuwen, et al.
Pubblicazione: (2024)
di: Zheng, Xiuwen, et al.
Pubblicazione: (2024)
The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge
di: Xue, Hongfei, et al.
Pubblicazione: (2025)
di: Xue, Hongfei, et al.
Pubblicazione: (2025)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
di: Huang, Mengcheng, et al.
Pubblicazione: (2026)
di: Huang, Mengcheng, et al.
Pubblicazione: (2026)
Sparsely Shared LoRA on Whisper for Child Speech Recognition
di: Liu, Wei, et al.
Pubblicazione: (2023)
di: Liu, Wei, et al.
Pubblicazione: (2023)
PCQ: Emotion Recognition in Speech via Progressive Channel Querying
di: Wang, Xincheng, et al.
Pubblicazione: (2024)
di: Wang, Xincheng, et al.
Pubblicazione: (2024)
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
di: Bondaruk, Łukasz, et al.
Pubblicazione: (2024)
di: Bondaruk, Łukasz, et al.
Pubblicazione: (2024)
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
di: Derington, Anna, et al.
Pubblicazione: (2023)
di: Derington, Anna, et al.
Pubblicazione: (2023)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
di: Farhadipour, Aref, et al.
Pubblicazione: (2024)
di: Farhadipour, Aref, et al.
Pubblicazione: (2024)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
di: Wang, Huimeng, et al.
Pubblicazione: (2025)
di: Wang, Huimeng, et al.
Pubblicazione: (2025)
Rehearsal-Free Online Continual Learning for Automatic Speech Recognition
di: Eeckt, Steven Vander, et al.
Pubblicazione: (2023)
di: Eeckt, Steven Vander, et al.
Pubblicazione: (2023)
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
di: Kutsakov, Aleksandr, et al.
Pubblicazione: (2025)
di: Kutsakov, Aleksandr, et al.
Pubblicazione: (2025)
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
di: Pusateri, Ernest, et al.
Pubblicazione: (2024)
di: Pusateri, Ernest, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
di: Nasretdinov, Rauf, et al.
Pubblicazione: (2025) -
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026) -
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
di: Alsayegh, Ali, et al.
Pubblicazione: (2025) -
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
di: Tian, Jingguang, et al.
Pubblicazione: (2024) -
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
di: Yen, Hao, et al.
Pubblicazione: (2024)