Saved in:
| Main Authors: | Wade, Papa Séga, Andries, Mihai, Kanellos, Ioannis, Moudenc, Thierry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.20243 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TIPAA-SSL: Text Independent Phone-to-Audio Alignment based on Self-Supervised Learning and Knowledge Transfer
by: Tits, Noé, et al.
Published: (2024)
by: Tits, Noé, et al.
Published: (2024)
Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies
by: Wu, Chung-Wen, et al.
Published: (2024)
by: Wu, Chung-Wen, et al.
Published: (2024)
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
by: Gulzar, Kashaf, et al.
Published: (2025)
by: Gulzar, Kashaf, et al.
Published: (2025)
Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
by: Gogoi, Parismita, et al.
Published: (2025)
by: Gogoi, Parismita, et al.
Published: (2025)
MMMOS: Multi-domain Multi-axis Audio Quality Assessment
by: Lin, Yi-Cheng, et al.
Published: (2025)
by: Lin, Yi-Cheng, et al.
Published: (2025)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
by: Prakash, Jeena, et al.
Published: (2025)
by: Prakash, Jeena, et al.
Published: (2025)
Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMs
by: Cuervo, Santiago, et al.
Published: (2025)
by: Cuervo, Santiago, et al.
Published: (2025)
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
by: Zhong, Jinzuomu, et al.
Published: (2023)
by: Zhong, Jinzuomu, et al.
Published: (2023)
Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
by: Frohmann, Markus, et al.
Published: (2025)
by: Frohmann, Markus, et al.
Published: (2025)
Handling Numeric Expressions in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Multimodal Proposal for an AI-Based Tool to Increase Cross-Assessment of Messages
by: Castro, Alejandro Álvarez, et al.
Published: (2025)
by: Castro, Alejandro Álvarez, et al.
Published: (2025)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
by: Zhang, Shucong, et al.
Published: (2025)
by: Zhang, Shucong, et al.
Published: (2025)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
by: He, Xinlu, et al.
Published: (2025)
by: He, Xinlu, et al.
Published: (2025)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
by: Min, Do June, et al.
Published: (2024)
by: Min, Do June, et al.
Published: (2024)
Optimizing Speech Multi-View Feature Fusion through Conditional Computation
by: Shan, Weiqiao, et al.
Published: (2025)
by: Shan, Weiqiao, et al.
Published: (2025)
Investigating Acoustic-Textual Emotional Inconsistency Information for Automatic Depression Detection
by: Su, Rongfeng, et al.
Published: (2024)
by: Su, Rongfeng, et al.
Published: (2024)
Unifying Model and Layer Fusion for Speech Foundation Models
by: Shih, Yi-Jen, et al.
Published: (2025)
by: Shih, Yi-Jen, et al.
Published: (2025)
Deep Learning for Assessment of Oral Reading Fluency
by: Vaidya, Mithilesh, et al.
Published: (2024)
by: Vaidya, Mithilesh, et al.
Published: (2024)
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain
by: Le-Duc, Khai
Published: (2024)
by: Le-Duc, Khai
Published: (2024)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
Gated Low-rank Adaptation for personalized Code-Switching Automatic Speech Recognition on the low-spec devices
by: Kim, Gwantae, et al.
Published: (2024)
by: Kim, Gwantae, et al.
Published: (2024)
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation
by: Khurana, Sameer, et al.
Published: (2023)
by: Khurana, Sameer, et al.
Published: (2023)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
by: Attia, Ahmed Adel, et al.
Published: (2024)
by: Attia, Ahmed Adel, et al.
Published: (2024)
Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
by: Parikh, Aditya Kamlesh, et al.
Published: (2026)
by: Parikh, Aditya Kamlesh, et al.
Published: (2026)
Benchmarking Automatic Speech Recognition for Indian Languages in Agricultural Contexts
by: S, Chandrashekar M, et al.
Published: (2026)
by: S, Chandrashekar M, et al.
Published: (2026)
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
by: Bannò, Stefano, et al.
Published: (2025)
by: Bannò, Stefano, et al.
Published: (2025)
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
by: Dai, Yusheng, et al.
Published: (2023)
by: Dai, Yusheng, et al.
Published: (2023)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
by: Nguyen, Tuan, et al.
Published: (2024)
by: Nguyen, Tuan, et al.
Published: (2024)
Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages
by: Pillai, Leena G, et al.
Published: (2024)
by: Pillai, Leena G, et al.
Published: (2024)
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring
by: Sudarshan, Ankitha, et al.
Published: (2023)
by: Sudarshan, Ankitha, et al.
Published: (2023)
Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
by: Ahasan, Md Mubtasim, et al.
Published: (2025)
by: Ahasan, Md Mubtasim, et al.
Published: (2025)
Incremental FastPitch: Chunk-based High Quality Text to Speech
by: Du, Muyang, et al.
Published: (2024)
by: Du, Muyang, et al.
Published: (2024)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
by: Yang, Cheng-Yeh, et al.
Published: (2026)
by: Yang, Cheng-Yeh, et al.
Published: (2026)
Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
by: Ling, Shaoshi, et al.
Published: (2025)
by: Ling, Shaoshi, et al.
Published: (2025)
Semantically Corrected Amharic Automatic Speech Recognition
by: Adnew, Samuael, et al.
Published: (2024)
by: Adnew, Samuael, et al.
Published: (2024)
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
by: Attia, Ahmed Adel, et al.
Published: (2025)
by: Attia, Ahmed Adel, et al.
Published: (2025)
CTC-Assisted LLM-Based Contextual ASR
by: Yang, Guanrou, et al.
Published: (2024)
by: Yang, Guanrou, et al.
Published: (2024)
Swin-BERT: A Feature Fusion System designed for Speech-based Alzheimer's Dementia Detection
by: Pan, Yilin, et al.
Published: (2024)
by: Pan, Yilin, et al.
Published: (2024)
Similar Items
-
TIPAA-SSL: Text Independent Phone-to-Audio Alignment based on Self-Supervised Learning and Knowledge Transfer
by: Tits, Noé, et al.
Published: (2024) -
Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies
by: Wu, Chung-Wen, et al.
Published: (2024) -
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
by: Gulzar, Kashaf, et al.
Published: (2025) -
Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
by: Gogoi, Parismita, et al.
Published: (2025) -
MMMOS: Multi-domain Multi-axis Audio Quality Assessment
by: Lin, Yi-Cheng, et al.
Published: (2025)