MedCutMix: A Data-Centric Approach to Improve Radiology Vision-Language Pre-training with Disease Awareness
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Sinuo, Xie, Yutong, Liu, Yuyuan, Wu, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-training Framework
by: Phan, Vu Minh Hieu, et al.
Published: (2024)
by: Phan, Vu Minh Hieu, et al.
Published: (2024)
PairAug: What Can Augmented Image-Text Pairs Do for Radiology?
by: Xie, Yutong, et al.
Published: (2024)
by: Xie, Yutong, et al.
Published: (2024)
VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training
by: Gong, Biao, et al.
Published: (2023)
by: Gong, Biao, et al.
Published: (2023)
Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training
by: Liang, Mingliang, et al.
Published: (2026)
by: Liang, Mingliang, et al.
Published: (2026)
A Reality Check of Vision-Language Pre-training in Radiology: Have We Progressed Using Text?
by: Silva-Rodríguez, Julio, et al.
Published: (2025)
by: Silva-Rodríguez, Julio, et al.
Published: (2025)
Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding
by: Shui, Zhongyi, et al.
Published: (2025)
by: Shui, Zhongyi, et al.
Published: (2025)
A Survey of Medical Vision-and-Language Applications and Their Techniques
by: Chen, Qi, et al.
Published: (2024)
by: Chen, Qi, et al.
Published: (2024)
MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset
by: Deng, Ziye, et al.
Published: (2025)
by: Deng, Ziye, et al.
Published: (2025)
Act Like a Radiologist: Radiology Report Generation across Anatomical Regions
by: Chen, Qi, et al.
Published: (2023)
by: Chen, Qi, et al.
Published: (2023)
Unsupervised Pre-training with Language-Vision Prompts for Low-Data Instance Segmentation
by: Zhang, Dingwen, et al.
Published: (2024)
by: Zhang, Dingwen, et al.
Published: (2024)
VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving
by: Zhang, Haiming, et al.
Published: (2024)
by: Zhang, Haiming, et al.
Published: (2024)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training
by: Wu, Biao, et al.
Published: (2024)
by: Wu, Biao, et al.
Published: (2024)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
by: Ma, Shuailei, et al.
Published: (2023)
by: Ma, Shuailei, et al.
Published: (2023)
Scaling Pre-training to One Hundred Billion Data for Vision Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
MedFILIP: Medical Fine-grained Language-Image Pre-training
by: Liang, Xinjie, et al.
Published: (2025)
by: Liang, Xinjie, et al.
Published: (2025)
3D Scene Graph Guided Vision-Language Pre-training
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Efficient Vision-Language Pre-training by Cluster Masking
by: Wei, Zihao, et al.
Published: (2024)
by: Wei, Zihao, et al.
Published: (2024)
Enhancing Vision-Language Pre-training with Rich Supervisions
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
Continual Retinal Vision-Language Pre-training upon Incremental Imaging Modalities
by: Yao, Yuang, et al.
Published: (2025)
by: Yao, Yuang, et al.
Published: (2025)
Intensive Vision-guided Network for Radiology Report Generation
by: Zheng, Fudan, et al.
Published: (2024)
by: Zheng, Fudan, et al.
Published: (2024)
Can Medical Vision-Language Pre-training Succeed with Purely Synthetic Data?
by: Liu, Che, et al.
Published: (2024)
by: Liu, Che, et al.
Published: (2024)
RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding
by: Fang, Tianchen, et al.
Published: (2025)
by: Fang, Tianchen, et al.
Published: (2025)
Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
by: Cao, Weiwei, et al.
Published: (2025)
by: Cao, Weiwei, et al.
Published: (2025)
Multilingual Vision-Language Pre-training for the Remote Sensing Domain
by: Silva, João Daniel, et al.
Published: (2024)
by: Silva, João Daniel, et al.
Published: (2024)
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
Unsupervised Domain Adaption Harnessing Vision-Language Pre-training
by: Zhou, Wenlve, et al.
Published: (2024)
by: Zhou, Wenlve, et al.
Published: (2024)
Data Augmentation in Human-Centric Vision
by: Jiang, Wentao, et al.
Published: (2024)
by: Jiang, Wentao, et al.
Published: (2024)
Omniview-Tuning: Boosting Viewpoint Invariance of Vision-Language Pre-training Models
by: Ruan, Shouwei, et al.
Published: (2024)
by: Ruan, Shouwei, et al.
Published: (2024)
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
by: Fhima, Jonathan, et al.
Published: (2024)
by: Fhima, Jonathan, et al.
Published: (2024)
RadCLIP: Enhancing Radiologic Image Analysis through Contrastive Language-Image Pre-training
by: Lu, Zhixiu, et al.
Published: (2024)
by: Lu, Zhixiu, et al.
Published: (2024)
NoiseCutMix: A Novel Data Augmentation Approach by Mixing Estimated Noise in Diffusion Models
by: Takezaki, Shumpei, et al.
Published: (2025)
by: Takezaki, Shumpei, et al.
Published: (2025)
Rethinking the Efficiency and Effectiveness of Reinforcement Learning for Radiology Report Generation
by: Lu, Zilin, et al.
Published: (2026)
by: Lu, Zilin, et al.
Published: (2026)
Conjugated Semantic Pool Improves OOD Detection with Pre-trained Vision-Language Models
by: Chen, Mengyuan, et al.
Published: (2024)
by: Chen, Mengyuan, et al.
Published: (2024)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
by: Wu, Shihan, et al.
Published: (2024)
by: Wu, Shihan, et al.
Published: (2024)
Integrating MedCLIP and Cross-Modal Fusion for Automatic Radiology Report Generation
by: Han, Qianhao, et al.
Published: (2024)
by: Han, Qianhao, et al.
Published: (2024)
Similar Items
-
Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-training Framework
by: Phan, Vu Minh Hieu, et al.
Published: (2024) -
PairAug: What Can Augmented Image-Text Pairs Do for Radiology?
by: Xie, Yutong, et al.
Published: (2024) -
VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
by: Zhang, Ziyang, et al.
Published: (2025) -
UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training
by: Gong, Biao, et al.
Published: (2023) -
Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training
by: Liang, Mingliang, et al.
Published: (2026)