MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908577378074624 |
|---|---|
| author | Kim, Soo Yong Cho, Suin Yun, Vincent-Daniel Hwang, Gyeongyeon |
| author_facet | Kim, Soo Yong Cho, Suin Yun, Vincent-Daniel Hwang, Gyeongyeon |
| contents | Bridging clinical diagnostic reasoning with AI remains a central challenge in medical imaging. We introduce MedCLM, an automated pipeline that converts detection datasets into large-scale medical visual question answering (VQA) data with Chain-of-Thought (CoT) reasoning by linking lesion boxes to organ segmentation and structured rationales. These contextual signals enable medical vision-language models to generate question-answer pairs with step-by-step reasoning. To utilize this data effectively, we propose an Integrated CoT-Curriculum Strategy composed of an Easy stage with explicit lesion boxes for visual grounding, a Medium stage that encourages implicit localization, and a Hard stage for weakly supervised reasoning. Experimental results demonstrate that MedCLM attains state-of-the-art performance on several medical VQA benchmarks, providing a scalable framework for developing clinically aligned medical vision-language models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_04477 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models Kim, Soo Yong Cho, Suin Yun, Vincent-Daniel Hwang, Gyeongyeon Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning Bridging clinical diagnostic reasoning with AI remains a central challenge in medical imaging. We introduce MedCLM, an automated pipeline that converts detection datasets into large-scale medical visual question answering (VQA) data with Chain-of-Thought (CoT) reasoning by linking lesion boxes to organ segmentation and structured rationales. These contextual signals enable medical vision-language models to generate question-answer pairs with step-by-step reasoning. To utilize this data effectively, we propose an Integrated CoT-Curriculum Strategy composed of an Easy stage with explicit lesion boxes for visual grounding, a Medium stage that encourages implicit localization, and a Hard stage for weakly supervised reasoning. Experimental results demonstrate that MedCLM attains state-of-the-art performance on several medical VQA benchmarks, providing a scalable framework for developing clinically aligned medical vision-language models. |
| title | MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2510.04477 |