Medical Spoken Named Entity Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Le-Duc, Khai, Thulke, David, Tran, Hung-Phong, Vo-Dang, Long, Nguyen, Khai-Nguyen, Hy, Truong-Son, Schlüter, Ralf |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Real-time Speech Summarization for Medical Conversations
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
Sentiment Reasoning for Healthcare
di: Nguyen, Khai-Nguyen, et al.
Pubblicazione: (2024)
di: Nguyen, Khai-Nguyen, et al.
Pubblicazione: (2024)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
di: Le-Duc, Khai, et al.
Pubblicazione: (2025)
di: Le-Duc, Khai, et al.
Pubblicazione: (2025)
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain
di: Le-Duc, Khai
Pubblicazione: (2024)
di: Le-Duc, Khai
Pubblicazione: (2024)
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
di: Pusateri, Ernest, et al.
Pubblicazione: (2024)
di: Pusateri, Ernest, et al.
Pubblicazione: (2024)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
di: Pham, The Hieu, et al.
Pubblicazione: (2025)
di: Pham, The Hieu, et al.
Pubblicazione: (2025)
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
di: Anh, Tran Nguyen, et al.
Pubblicazione: (2025)
di: Anh, Tran Nguyen, et al.
Pubblicazione: (2025)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
di: Le, Khanh, et al.
Pubblicazione: (2025)
di: Le, Khanh, et al.
Pubblicazione: (2025)
Revisiting Deep Audio-Text Retrieval Through the Lens of Transportation
di: Luong, Manh, et al.
Pubblicazione: (2024)
di: Luong, Manh, et al.
Pubblicazione: (2024)
AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
di: Nguyen-Le, Hai-Son, et al.
Pubblicazione: (2026)
di: Nguyen-Le, Hai-Son, et al.
Pubblicazione: (2026)
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
di: Le, Khanh, et al.
Pubblicazione: (2025)
di: Le, Khanh, et al.
Pubblicazione: (2025)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
di: Doan, Khai Duy, et al.
Pubblicazione: (2024)
di: Doan, Khai Duy, et al.
Pubblicazione: (2024)
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
di: Wang, Peng, et al.
Pubblicazione: (2023)
di: Wang, Peng, et al.
Pubblicazione: (2023)
Continuous Learning of Transformer-based Audio Deepfake Detection
di: Le, Tuan Duy Nguyen, et al.
Pubblicazione: (2024)
di: Le, Tuan Duy Nguyen, et al.
Pubblicazione: (2024)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
di: Raissi, Tina, et al.
Pubblicazione: (2025)
di: Raissi, Tina, et al.
Pubblicazione: (2025)
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization
di: Ho, Luong, et al.
Pubblicazione: (2025)
di: Ho, Luong, et al.
Pubblicazione: (2025)
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
di: Tu, Huu Tuong, et al.
Pubblicazione: (2025)
di: Tu, Huu Tuong, et al.
Pubblicazione: (2025)
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
di: Sasu, David, et al.
Pubblicazione: (2025)
di: Sasu, David, et al.
Pubblicazione: (2025)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
di: Le, Khanh, et al.
Pubblicazione: (2025)
di: Le, Khanh, et al.
Pubblicazione: (2025)
CLAP-Based Automatic Word Naming Recognition in Post-Stroke Aphasia
di: Kaloga, Yacouba, et al.
Pubblicazione: (2026)
di: Kaloga, Yacouba, et al.
Pubblicazione: (2026)
Comparison Performance of Spectrogram and Scalogram as Input of Acoustic Recognition Task
di: Phan, Dang Thoai
Pubblicazione: (2024)
di: Phan, Dang Thoai
Pubblicazione: (2024)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
di: Zeineldeen, Mohammad, et al.
Pubblicazione: (2023)
di: Zeineldeen, Mohammad, et al.
Pubblicazione: (2023)
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
di: Yang, Zijian, et al.
Pubblicazione: (2026)
di: Yang, Zijian, et al.
Pubblicazione: (2026)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
di: Dang, Trung, et al.
Pubblicazione: (2024)
di: Dang, Trung, et al.
Pubblicazione: (2024)
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
di: Le, Trang, et al.
Pubblicazione: (2024)
di: Le, Trang, et al.
Pubblicazione: (2024)
Zero-Shot Text-to-Speech from Continuous Text Streams
di: Dang, Trung, et al.
Pubblicazione: (2024)
di: Dang, Trung, et al.
Pubblicazione: (2024)
Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
di: Raissi, Tina, et al.
Pubblicazione: (2024)
di: Raissi, Tina, et al.
Pubblicazione: (2024)
Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
di: Hoang, Long-Vu, et al.
Pubblicazione: (2025)
di: Hoang, Long-Vu, et al.
Pubblicazione: (2025)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
di: Vo, Quoc Thinh, et al.
Pubblicazione: (2025)
di: Vo, Quoc Thinh, et al.
Pubblicazione: (2025)
Room Impulse Responses help attackers to evade Deep Fake Detection
di: Luong, Hieu-Thi, et al.
Pubblicazione: (2024)
di: Luong, Hieu-Thi, et al.
Pubblicazione: (2024)
On the use of Performer and Agent Attention for Spoken Language Identification
di: dhiman, Jitendra Kumar, et al.
Pubblicazione: (2025)
di: dhiman, Jitendra Kumar, et al.
Pubblicazione: (2025)
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
di: Vu, Hoang Long, et al.
Pubblicazione: (2024)
di: Vu, Hoang Long, et al.
Pubblicazione: (2024)
Xi+: Uncertainty Supervision for Robust Speaker Embedding
di: Li, Junjie, et al.
Pubblicazione: (2025)
di: Li, Junjie, et al.
Pubblicazione: (2025)
Keyword Mamba: Spoken Keyword Spotting with State Space Models
di: Ding, Hanyu, et al.
Pubblicazione: (2025)
di: Ding, Hanyu, et al.
Pubblicazione: (2025)
Spoken language change detection inspired by speaker change detection
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Real-time Speech Summarization for Medical Conversations
di: Le-Duc, Khai, et al.
Pubblicazione: (2024) -
Sentiment Reasoning for Healthcare
di: Nguyen, Khai-Nguyen, et al.
Pubblicazione: (2024) -
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
di: Le-Duc, Khai, et al.
Pubblicazione: (2024) -
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
di: Le-Duc, Khai, et al.
Pubblicazione: (2024) -
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)