VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910982511525888 |
|---|---|
| author | Mehta, Manas Pan, Yimu Gallagher, Kelly Gernand, Alison D. Goldstein, Jeffery A. Mwinyelle, Delia Mithal, Leena Wang, James Z. |
| author_facet | Mehta, Manas Pan, Yimu Gallagher, Kelly Gernand, Alison D. Goldstein, Jeffery A. Mwinyelle, Delia Mithal, Leena Wang, James Z. |
| contents | Pathological examination of the placenta is an effective method for detecting and mitigating health risks associated with childbirth. Recent advancements in AI have enabled the use of photographs of the placenta and pathology reports for detecting and classifying signs of childbirth-related pathologies. However, existing automated methods are computationally extensive, which limits their deployability. We propose two modifications to vision-language contrastive learning (VLC) frameworks to enhance their accuracy and efficiency: (1) text-anchored vision-language contrastive knowledge distillation (VLCD)-a new knowledge distillation strategy for medical VLC pretraining, and (2) unsupervised predistillation using a large natural images dataset for improved initialization. Our approach distills efficient neural networks that match or surpass the teacher model in performance while achieving model compression and acceleration. Our results showcase the value of unsupervised predistillation in improving the performance and robustness of our approach, specifically for lower-quality images. VLCD serves as an effective way to improve the efficiency and deployability of medical VLC approaches, making AI-based healthcare solutions more accessible, especially in resource-constrained environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_02229 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis Mehta, Manas Pan, Yimu Gallagher, Kelly Gernand, Alison D. Goldstein, Jeffery A. Mwinyelle, Delia Mithal, Leena Wang, James Z. Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning Pathological examination of the placenta is an effective method for detecting and mitigating health risks associated with childbirth. Recent advancements in AI have enabled the use of photographs of the placenta and pathology reports for detecting and classifying signs of childbirth-related pathologies. However, existing automated methods are computationally extensive, which limits their deployability. We propose two modifications to vision-language contrastive learning (VLC) frameworks to enhance their accuracy and efficiency: (1) text-anchored vision-language contrastive knowledge distillation (VLCD)-a new knowledge distillation strategy for medical VLC pretraining, and (2) unsupervised predistillation using a large natural images dataset for improved initialization. Our approach distills efficient neural networks that match or surpass the teacher model in performance while achieving model compression and acceleration. Our results showcase the value of unsupervised predistillation in improving the performance and robustness of our approach, specifically for lower-quality images. VLCD serves as an effective way to improve the efficiency and deployability of medical VLC approaches, making AI-based healthcare solutions more accessible, especially in resource-constrained environments. |
| title | VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2506.02229 |