CBF-AFA: Chunk-Based Multi-SSL Fusion for Automatic Fluency Assessment
Fuente:
arXiv
Guardado en:
| Autores principales: | Wade, Papa Séga, Andries, Mihai, Kanellos, Ioannis, Moudenc, Thierry |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TIPAA-SSL: Text Independent Phone-to-Audio Alignment based on Self-Supervised Learning and Knowledge Transfer
por: Tits, Noé, et al.
Publicado: (2024)
por: Tits, Noé, et al.
Publicado: (2024)
Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies
por: Wu, Chung-Wen, et al.
Publicado: (2024)
por: Wu, Chung-Wen, et al.
Publicado: (2024)
MMMOS: Multi-domain Multi-axis Audio Quality Assessment
por: Lin, Yi-Cheng, et al.
Publicado: (2025)
por: Lin, Yi-Cheng, et al.
Publicado: (2025)
Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
por: Gogoi, Parismita, et al.
Publicado: (2025)
por: Gogoi, Parismita, et al.
Publicado: (2025)
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
por: Gulzar, Kashaf, et al.
Publicado: (2025)
por: Gulzar, Kashaf, et al.
Publicado: (2025)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
por: Prakash, Jeena, et al.
Publicado: (2025)
por: Prakash, Jeena, et al.
Publicado: (2025)
Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMs
por: Cuervo, Santiago, et al.
Publicado: (2025)
por: Cuervo, Santiago, et al.
Publicado: (2025)
Handling Numeric Expressions in Automatic Speech Recognition
por: Huber, Christian, et al.
Publicado: (2024)
por: Huber, Christian, et al.
Publicado: (2024)
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
por: Zhong, Jinzuomu, et al.
Publicado: (2023)
por: Zhong, Jinzuomu, et al.
Publicado: (2023)
Multimodal Proposal for an AI-Based Tool to Increase Cross-Assessment of Messages
por: Castro, Alejandro Álvarez, et al.
Publicado: (2025)
por: Castro, Alejandro Álvarez, et al.
Publicado: (2025)
Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
por: Frohmann, Markus, et al.
Publicado: (2025)
por: Frohmann, Markus, et al.
Publicado: (2025)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
por: Zhang, Shucong, et al.
Publicado: (2025)
por: Zhang, Shucong, et al.
Publicado: (2025)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
por: Min, Do June, et al.
Publicado: (2024)
por: Min, Do June, et al.
Publicado: (2024)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
por: He, Xinlu, et al.
Publicado: (2025)
por: He, Xinlu, et al.
Publicado: (2025)
Investigating Acoustic-Textual Emotional Inconsistency Information for Automatic Depression Detection
por: Su, Rongfeng, et al.
Publicado: (2024)
por: Su, Rongfeng, et al.
Publicado: (2024)
Unifying Model and Layer Fusion for Speech Foundation Models
por: Shih, Yi-Jen, et al.
Publicado: (2025)
por: Shih, Yi-Jen, et al.
Publicado: (2025)
Optimizing Speech Multi-View Feature Fusion through Conditional Computation
por: Shan, Weiqiao, et al.
Publicado: (2025)
por: Shan, Weiqiao, et al.
Publicado: (2025)
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain
por: Le-Duc, Khai
Publicado: (2024)
por: Le-Duc, Khai
Publicado: (2024)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
Deep Learning for Assessment of Oral Reading Fluency
por: Vaidya, Mithilesh, et al.
Publicado: (2024)
por: Vaidya, Mithilesh, et al.
Publicado: (2024)
Gated Low-rank Adaptation for personalized Code-Switching Automatic Speech Recognition on the low-spec devices
por: Kim, Gwantae, et al.
Publicado: (2024)
por: Kim, Gwantae, et al.
Publicado: (2024)
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
por: Do, Heejin, et al.
Publicado: (2024)
por: Do, Heejin, et al.
Publicado: (2024)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
por: Attia, Ahmed Adel, et al.
Publicado: (2024)
por: Attia, Ahmed Adel, et al.
Publicado: (2024)
Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
por: Parikh, Aditya Kamlesh, et al.
Publicado: (2026)
por: Parikh, Aditya Kamlesh, et al.
Publicado: (2026)
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation
por: Khurana, Sameer, et al.
Publicado: (2023)
por: Khurana, Sameer, et al.
Publicado: (2023)
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
por: Bannò, Stefano, et al.
Publicado: (2025)
por: Bannò, Stefano, et al.
Publicado: (2025)
Benchmarking Automatic Speech Recognition for Indian Languages in Agricultural Contexts
por: S, Chandrashekar M, et al.
Publicado: (2026)
por: S, Chandrashekar M, et al.
Publicado: (2026)
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
por: Dai, Yusheng, et al.
Publicado: (2023)
por: Dai, Yusheng, et al.
Publicado: (2023)
Incremental FastPitch: Chunk-based High Quality Text to Speech
por: Du, Muyang, et al.
Publicado: (2024)
por: Du, Muyang, et al.
Publicado: (2024)
Zero-Shot Text-to-Speech as Golden Speech Generator: A Systematic Framework and its Applicability in Automatic Pronunciation Assessment
por: Lo, Tien-Hong, et al.
Publicado: (2024)
por: Lo, Tien-Hong, et al.
Publicado: (2024)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
por: Nguyen, Tuan, et al.
Publicado: (2024)
por: Nguyen, Tuan, et al.
Publicado: (2024)
Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages
por: Pillai, Leena G, et al.
Publicado: (2024)
por: Pillai, Leena G, et al.
Publicado: (2024)
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring
por: Sudarshan, Ankitha, et al.
Publicado: (2023)
por: Sudarshan, Ankitha, et al.
Publicado: (2023)
Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
por: Zhang, Wei, et al.
Publicado: (2025)
por: Zhang, Wei, et al.
Publicado: (2025)
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
por: Ahasan, Md Mubtasim, et al.
Publicado: (2025)
por: Ahasan, Md Mubtasim, et al.
Publicado: (2025)
Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
por: Ling, Shaoshi, et al.
Publicado: (2025)
por: Ling, Shaoshi, et al.
Publicado: (2025)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
por: Lau, Hok-Shing, et al.
Publicado: (2024)
por: Lau, Hok-Shing, et al.
Publicado: (2024)
CTC-Assisted LLM-Based Contextual ASR
por: Yang, Guanrou, et al.
Publicado: (2024)
por: Yang, Guanrou, et al.
Publicado: (2024)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
por: Yang, Cheng-Yeh, et al.
Publicado: (2026)
por: Yang, Cheng-Yeh, et al.
Publicado: (2026)
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
por: Wang, Qiongqiong, et al.
Publicado: (2025)
por: Wang, Qiongqiong, et al.
Publicado: (2025)
Ejemplares similares
-
TIPAA-SSL: Text Independent Phone-to-Audio Alignment based on Self-Supervised Learning and Knowledge Transfer
por: Tits, Noé, et al.
Publicado: (2024) -
Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies
por: Wu, Chung-Wen, et al.
Publicado: (2024) -
MMMOS: Multi-domain Multi-axis Audio Quality Assessment
por: Lin, Yi-Cheng, et al.
Publicado: (2025) -
Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
por: Gogoi, Parismita, et al.
Publicado: (2025) -
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
por: Gulzar, Kashaf, et al.
Publicado: (2025)