Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qian, Mengjie, Ma, Rao, Bannò, Stefano, Knill, Kate M., Gales, Mark J. F. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
Towards End-to-End Spoken Grammatical Error Correction
von: Bannò, Stefano, et al.
Veröffentlicht: (2023)
von: Bannò, Stefano, et al.
Veröffentlicht: (2023)
Data Augmentation for Spoken Grammatical Error Correction
von: Karanasou, Penny, et al.
Veröffentlicht: (2025)
von: Karanasou, Penny, et al.
Veröffentlicht: (2025)
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024)
von: Ma, Rao, et al.
Veröffentlicht: (2024)
Assessment of L2 Oral Proficiency using Speech Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness
von: Loweimi, Erfan, et al.
Veröffentlicht: (2025)
von: Loweimi, Erfan, et al.
Veröffentlicht: (2025)
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
von: Le, Trang, et al.
Veröffentlicht: (2024)
von: Le, Trang, et al.
Veröffentlicht: (2024)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
Training Articulatory Inversion Models for Interspeaker Consistency
von: McGhee, Charles, et al.
Veröffentlicht: (2025)
von: McGhee, Charles, et al.
Veröffentlicht: (2025)
Retrieval Augmented End-to-End Spoken Dialog Models
von: Wang, Mingqiu, et al.
Veröffentlicht: (2024)
von: Wang, Mingqiu, et al.
Veröffentlicht: (2024)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
von: Lin, Chyi-Jiunn, et al.
Veröffentlicht: (2024)
von: Lin, Chyi-Jiunn, et al.
Veröffentlicht: (2024)
FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning
von: Chen, Tanyu, et al.
Veröffentlicht: (2026)
von: Chen, Tanyu, et al.
Veröffentlicht: (2026)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
An End-to-End Speech Summarization Using Large Language Model
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
von: Prabhavalkar, Rohit, et al.
Veröffentlicht: (2024)
von: Prabhavalkar, Rohit, et al.
Veröffentlicht: (2024)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
Alternating Weak Triphone/BPE Alignment Supervision from Hybrid Model Improves End-to-End ASR
von: Jiang, Jintao, et al.
Veröffentlicht: (2024)
von: Jiang, Jintao, et al.
Veröffentlicht: (2024)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
End-to-End Speech-to-Text Translation: A Survey
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
von: Li, Chia-Yu, et al.
Veröffentlicht: (2024)
von: Li, Chia-Yu, et al.
Veröffentlicht: (2024)
An investigation of phrase break prediction in an End-to-End TTS system
von: Vadapalli, Anandaswarup
Veröffentlicht: (2023)
von: Vadapalli, Anandaswarup
Veröffentlicht: (2023)
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
von: Wang, Peidong, et al.
Veröffentlicht: (2026)
von: Wang, Peidong, et al.
Veröffentlicht: (2026)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
Evolutionary Prompt Design for LLM-Based Post-ASR Error Correction
von: Sachdev, Rithik, et al.
Veröffentlicht: (2024)
von: Sachdev, Rithik, et al.
Veröffentlicht: (2024)
Acoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems
von: Lin, Jhen-Ke, et al.
Veröffentlicht: (2025)
von: Lin, Jhen-Ke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025) -
Towards End-to-End Spoken Grammatical Error Correction
von: Bannò, Stefano, et al.
Veröffentlicht: (2023) -
Data Augmentation for Spoken Grammatical Error Correction
von: Karanasou, Penny, et al.
Veröffentlicht: (2025) -
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024) -
Assessment of L2 Oral Proficiency using Speech Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2025)