VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain
Fuente:
arXiv
Salvato in:
| Autore principale: | Le-Duc, Khai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Real-time Speech Summarization for Medical Conversations
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
PhoWhisper: Automatic Speech Recognition for Vietnamese
di: Le, Thanh-Thien, et al.
Pubblicazione: (2024)
di: Le, Thanh-Thien, et al.
Pubblicazione: (2024)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
di: Zhang, Shucong, et al.
Pubblicazione: (2025)
di: Zhang, Shucong, et al.
Pubblicazione: (2025)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
di: Zhuo, Jianheng, et al.
Pubblicazione: (2025)
di: Zhuo, Jianheng, et al.
Pubblicazione: (2025)
Benchmarking Automatic Speech Recognition for Indian Languages in Agricultural Contexts
di: S, Chandrashekar M, et al.
Pubblicazione: (2026)
di: S, Chandrashekar M, et al.
Pubblicazione: (2026)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
di: Min, Do June, et al.
Pubblicazione: (2024)
di: Min, Do June, et al.
Pubblicazione: (2024)
Handling Numeric Expressions in Automatic Speech Recognition
di: Huber, Christian, et al.
Pubblicazione: (2024)
di: Huber, Christian, et al.
Pubblicazione: (2024)
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
di: Le-Duc, Khai, et al.
Pubblicazione: (2025)
di: Le-Duc, Khai, et al.
Pubblicazione: (2025)
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
di: Anh, Tran Nguyen, et al.
Pubblicazione: (2025)
di: Anh, Tran Nguyen, et al.
Pubblicazione: (2025)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring
di: Sudarshan, Ankitha, et al.
Pubblicazione: (2023)
di: Sudarshan, Ankitha, et al.
Pubblicazione: (2023)
Gated Low-rank Adaptation for personalized Code-Switching Automatic Speech Recognition on the low-spec devices
di: Kim, Gwantae, et al.
Pubblicazione: (2024)
di: Kim, Gwantae, et al.
Pubblicazione: (2024)
Semantically Corrected Amharic Automatic Speech Recognition
di: Adnew, Samuael, et al.
Pubblicazione: (2024)
di: Adnew, Samuael, et al.
Pubblicazione: (2024)
Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages
di: Pillai, Leena G, et al.
Pubblicazione: (2024)
di: Pillai, Leena G, et al.
Pubblicazione: (2024)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
di: He, Xinlu, et al.
Pubblicazione: (2025)
di: He, Xinlu, et al.
Pubblicazione: (2025)
Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
di: Zhang, Wei, et al.
Pubblicazione: (2025)
di: Zhang, Wei, et al.
Pubblicazione: (2025)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
di: Wang, Chien-Chun, et al.
Pubblicazione: (2024)
di: Wang, Chien-Chun, et al.
Pubblicazione: (2024)
Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
di: Gedeon, Máté, et al.
Pubblicazione: (2025)
di: Gedeon, Máté, et al.
Pubblicazione: (2025)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2026)
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2026)
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects
di: Yang, Sicheng, et al.
Pubblicazione: (2026)
di: Yang, Sicheng, et al.
Pubblicazione: (2026)
Sentiment Reasoning for Healthcare
di: Nguyen, Khai-Nguyen, et al.
Pubblicazione: (2024)
di: Nguyen, Khai-Nguyen, et al.
Pubblicazione: (2024)
Towards Robust Speech Recognition for Jamaican Patois Music Transcription
di: Madden, Jordan, et al.
Pubblicazione: (2025)
di: Madden, Jordan, et al.
Pubblicazione: (2025)
CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
di: Carvalho, Carlos, et al.
Pubblicazione: (2025)
di: Carvalho, Carlos, et al.
Pubblicazione: (2025)
Enhancing ASR Performance in the Medical Domain for Dravidian Languages
di: Devarakonda, Sri Charan, et al.
Pubblicazione: (2026)
di: Devarakonda, Sri Charan, et al.
Pubblicazione: (2026)
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge
di: Li, Xiaoxiao, et al.
Pubblicazione: (2025)
di: Li, Xiaoxiao, et al.
Pubblicazione: (2025)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
di: Storey, Edward, et al.
Pubblicazione: (2025)
di: Storey, Edward, et al.
Pubblicazione: (2025)
Medical Spoken Named Entity Recognition
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
A Large Dataset of Spontaneous Speech with the Accent Spoken in São Paulo for Automatic Speech Recognition Evaluation
di: Lima, Rodrigo, et al.
Pubblicazione: (2024)
di: Lima, Rodrigo, et al.
Pubblicazione: (2024)
SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
di: Pandey, Prabhat, et al.
Pubblicazione: (2025)
di: Pandey, Prabhat, et al.
Pubblicazione: (2025)
Reading Miscue Detection in Primary School through Automatic Speech Recognition
di: Gao, Lingyun, et al.
Pubblicazione: (2024)
di: Gao, Lingyun, et al.
Pubblicazione: (2024)
Efficient Streaming LLM for Speech Recognition
di: Jia, Junteng, et al.
Pubblicazione: (2024)
di: Jia, Junteng, et al.
Pubblicazione: (2024)
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
di: Moure, Pehuén, et al.
Pubblicazione: (2026)
di: Moure, Pehuén, et al.
Pubblicazione: (2026)
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation
di: Khurana, Sameer, et al.
Pubblicazione: (2023)
di: Khurana, Sameer, et al.
Pubblicazione: (2023)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Real-time Speech Summarization for Medical Conversations
di: Le-Duc, Khai, et al.
Pubblicazione: (2024) -
PhoWhisper: Automatic Speech Recognition for Vietnamese
di: Le, Thanh-Thien, et al.
Pubblicazione: (2024) -
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
di: Zhang, Shucong, et al.
Pubblicazione: (2025) -
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
di: Le-Duc, Khai, et al.
Pubblicazione: (2024) -
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
di: Zhuo, Jianheng, et al.
Pubblicazione: (2025)