MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914310569066496 |
|---|---|
| author | Wang, Wenjie Wu, Wei Liu, Ying Zhao, Yuan Lv, Xiaole Diao, Liang Fan, Zengjian Xie, Wenfeng Lin, Ziling Shi, De Huang, Lin Xu, Kaihe Li, Hong |
| author_facet | Wang, Wenjie Wu, Wei Liu, Ying Zhao, Yuan Lv, Xiaole Diao, Liang Fan, Zengjian Xie, Wenfeng Lin, Ziling Shi, De Huang, Lin Xu, Kaihe Li, Hong |
| contents | Medical document OCR is challenging due to complex layouts, domain-specific terminology, and noisy annotations, while requiring strict field-level exact matching. Existing OCR systems and general-purpose vision-language models often fail to reliably parse such documents. We propose MeDocVL, a post-trained vision-language model for query-driven medical document parsing. Our framework combines Training-driven Label Refinement to construct high-quality supervision from noisy annotations, with a Noise-aware Hybrid Post-training strategy that integrates reinforcement learning and supervised fine-tuning to achieve robust and precise extraction. Experiments on medical invoice benchmarks show that MeDocVL consistently outperforms conventional OCR systems and strong VLM baselines, achieving state-of-the-art performance under noisy supervision. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_06402 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing Wang, Wenjie Wu, Wei Liu, Ying Zhao, Yuan Lv, Xiaole Diao, Liang Fan, Zengjian Xie, Wenfeng Lin, Ziling Shi, De Huang, Lin Xu, Kaihe Li, Hong Computer Vision and Pattern Recognition Medical document OCR is challenging due to complex layouts, domain-specific terminology, and noisy annotations, while requiring strict field-level exact matching. Existing OCR systems and general-purpose vision-language models often fail to reliably parse such documents. We propose MeDocVL, a post-trained vision-language model for query-driven medical document parsing. Our framework combines Training-driven Label Refinement to construct high-quality supervision from noisy annotations, with a Noise-aware Hybrid Post-training strategy that integrates reinforcement learning and supervised fine-tuning to achieve robust and precise extraction. Experiments on medical invoice benchmarks show that MeDocVL consistently outperforms conventional OCR systems and strong VLM baselines, achieving state-of-the-art performance under noisy supervision. |
| title | MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.06402 |