Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Parikh, Aditya Kamlesh, Tejedor-Garcia, Cristian, Cucchiarini, Catia, Strik, Helmer |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-Shot Speech LLMs for Multi-Aspect Evaluation of L2 Speech: Challenges and Opportunities
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026)
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026)
Evaluating Logit-Based GOP Scores for Mispronunciation Detection
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2025)
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2025)
Improving Child Speech Recognition and Reading Mistake Detection by Using Prompts
von: Gao, Lingyun, et al.
Veröffentlicht: (2025)
von: Gao, Lingyun, et al.
Veröffentlicht: (2025)
Automatic Assessment of Oral Reading Accuracy for Reading Diagnostics
von: Molenaar, Bo, et al.
Veröffentlicht: (2023)
von: Molenaar, Bo, et al.
Veröffentlicht: (2023)
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
von: Wills, Simone, et al.
Veröffentlicht: (2023)
von: Wills, Simone, et al.
Veröffentlicht: (2023)
Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological Knowledge
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2025)
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2025)
An ASR-Based Tutor for Learning to Read: How to Optimize Feedback to First Graders
von: Bai, Yu, et al.
Veröffentlicht: (2023)
von: Bai, Yu, et al.
Veröffentlicht: (2023)
Reading Miscue Detection in Primary School through Automatic Speech Recognition
von: Gao, Lingyun, et al.
Veröffentlicht: (2024)
von: Gao, Lingyun, et al.
Veröffentlicht: (2024)
Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech
von: Lathouwers, Gus, et al.
Veröffentlicht: (2026)
von: Lathouwers, Gus, et al.
Veröffentlicht: (2026)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models
von: Ding, Shaojin, et al.
Veröffentlicht: (2023)
von: Ding, Shaojin, et al.
Veröffentlicht: (2023)
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion
von: Geng, Haopeng, et al.
Veröffentlicht: (2025)
von: Geng, Haopeng, et al.
Veröffentlicht: (2025)
Attention-Guided Adaptation for Code-Switching Speech Recognition
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and Pauses
von: Gómez-Zaragozá, Lucía, et al.
Veröffentlicht: (2023)
von: Gómez-Zaragozá, Lucía, et al.
Veröffentlicht: (2023)
LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
von: Yemini, Yochai, et al.
Veröffentlicht: (2023)
von: Yemini, Yochai, et al.
Veröffentlicht: (2023)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
von: Wang, Wei, et al.
Veröffentlicht: (2025)
von: Wang, Wei, et al.
Veröffentlicht: (2025)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
von: Geng, Haopeng, et al.
Veröffentlicht: (2024)
von: Geng, Haopeng, et al.
Veröffentlicht: (2024)
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
Fine-Grained and Interpretable Neural Speech Editing
von: Morrison, Max, et al.
Veröffentlicht: (2024)
von: Morrison, Max, et al.
Veröffentlicht: (2024)
Multi-modal Speech Enhancement with Limited Electromyography Channels
von: Feng, Fuyuan, et al.
Veröffentlicht: (2025)
von: Feng, Fuyuan, et al.
Veröffentlicht: (2025)
Unsupervised Multi-channel Speech Dereverberation via Diffusion
von: Wu, Yulun, et al.
Veröffentlicht: (2025)
von: Wu, Yulun, et al.
Veröffentlicht: (2025)
Attention-Based Beamformer For Multi-Channel Speech Enhancement
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Zero-Shot Speech LLMs for Multi-Aspect Evaluation of L2 Speech: Challenges and Opportunities
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026) -
Evaluating Logit-Based GOP Scores for Mispronunciation Detection
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2025) -
Improving Child Speech Recognition and Reading Mistake Detection by Using Prompts
von: Gao, Lingyun, et al.
Veröffentlicht: (2025) -
Automatic Assessment of Oral Reading Accuracy for Reading Diagnostics
von: Molenaar, Bo, et al.
Veröffentlicht: (2023) -
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
von: Wills, Simone, et al.
Veröffentlicht: (2023)