Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
Fuente:
arXiv
Saved in:
| Main Authors: | Pham, Phuc, Pham, Nhu, Ly, Ngoc Quoc |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Recent Advances in Medical Image Classification
by: Dao, Loan, et al.
Published: (2025)
by: Dao, Loan, et al.
Published: (2025)
A Comprehensive Study on Medical Image Segmentation using Deep Neural Networks
by: Dao, Loan, et al.
Published: (2025)
by: Dao, Loan, et al.
Published: (2025)
Enhancing Feature Diversity Boosts Channel-Adaptive Vision Transformers
by: Pham, Chau, et al.
Published: (2024)
by: Pham, Chau, et al.
Published: (2024)
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
by: Tran, Quoc-Khang, et al.
Published: (2026)
by: Tran, Quoc-Khang, et al.
Published: (2026)
Ontology-based knowledge representation for bone disease diagnosis: a foundation for safe and sustainable medical artificial intelligence systems
by: Dao, Loan, et al.
Published: (2025)
by: Dao, Loan, et al.
Published: (2025)
Enhancing Medical Large Vision-Language Models via Alignment Distillation
by: Chang, Aofei, et al.
Published: (2025)
by: Chang, Aofei, et al.
Published: (2025)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging
by: Kumar, Ashwin, et al.
Published: (2026)
by: Kumar, Ashwin, et al.
Published: (2026)
Multi-Aspect Knowledge-Enhanced Medical Vision-Language Pretraining with Multi-Agent Data Generation
by: Li, Xieji, et al.
Published: (2025)
by: Li, Xieji, et al.
Published: (2025)
Boost Self-Supervised Dataset Distillation via Parameterization, Predefined Augmentation, and Approximation
by: Yu, Sheng-Feng, et al.
Published: (2025)
by: Yu, Sheng-Feng, et al.
Published: (2025)
Aleatoric Uncertainty Medical Image Segmentation Estimation via Flow Matching
by: Van Nguyen, Phi, et al.
Published: (2025)
by: Van Nguyen, Phi, et al.
Published: (2025)
ALPI: Auto-Labeller with Proxy Injection for 3D Object Detection using 2D Labels Only
by: Lahlali, Saad, et al.
Published: (2024)
by: Lahlali, Saad, et al.
Published: (2024)
Diffusion Model in Latent Space for Medical Image Segmentation Task
by: Ngoc, Huynh Trinh, et al.
Published: (2025)
by: Ngoc, Huynh Trinh, et al.
Published: (2025)
OE3DIS: Open-Ended 3D Point Cloud Instance Segmentation
by: Nguyen, Phuc D. A., et al.
Published: (2024)
by: Nguyen, Phuc D. A., et al.
Published: (2024)
TinySSL: Distilled Self-Supervised Pretraining for Sub-Megabyte MCU Models
by: Wilson, Bibin
Published: (2026)
by: Wilson, Bibin
Published: (2026)
QTSeg: A Query Token-Based Dual-Mix Attention Framework with Multi-Level Feature Distribution for Medical Image Segmentation
by: Tran, Phuong-Nam, et al.
Published: (2024)
by: Tran, Phuong-Nam, et al.
Published: (2024)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
by: Vu, Sinh Trong, et al.
Published: (2025)
by: Vu, Sinh Trong, et al.
Published: (2025)
Confounder-Aware Medical Data Selection for Fine-Tuning Pretrained Vision Models
by: Ji, Anyang, et al.
Published: (2025)
by: Ji, Anyang, et al.
Published: (2025)
Comparative Study of UNet-based Architectures for Liver Tumor Segmentation in Multi-Phase Contrast-Enhanced Computed Tomography
by: Ly, Doan-Van-Anh, et al.
Published: (2025)
by: Ly, Doan-Van-Anh, et al.
Published: (2025)
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
by: Quang, Ngoc Bui Lam, et al.
Published: (2025)
by: Quang, Ngoc Bui Lam, et al.
Published: (2025)
SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment
by: Lu, Wenbo
Published: (2025)
by: Lu, Wenbo
Published: (2025)
DINORANKCLIP: DINOv3 Distillation and Injection for Vision-Language Pretraining with High-Order Ranking Consistency
by: Jiang, Shuyang, et al.
Published: (2026)
by: Jiang, Shuyang, et al.
Published: (2026)
LiteNeXt: A Novel Lightweight ConvMixer-based Model with Self-embedding Representation Parallel for Medical Image Segmentation
by: Tran, Ngoc-Du, et al.
Published: (2024)
by: Tran, Ngoc-Du, et al.
Published: (2024)
UMSPU: Universal Multi-Size Phase Unwrapping via Mutual Self-Distillation and Adaptive Boosting Ensemble Segmenters
by: Du, Lintong, et al.
Published: (2024)
by: Du, Lintong, et al.
Published: (2024)
HDC: Hierarchical Distillation for Multi-level Noisy Consistency in Semi-Supervised Fetal Ultrasound Segmentation
by: Le, Tran Quoc Khanh, et al.
Published: (2025)
by: Le, Tran Quoc Khanh, et al.
Published: (2025)
ThyroidEffi 1.0: A Cost-Effective System for High-Performance Multi-Class Thyroid Carcinoma Classification
by: Pham-Ngoc, Hai, et al.
Published: (2025)
by: Pham-Ngoc, Hai, et al.
Published: (2025)
Person Re-Identification System at Semantic Level based on Pedestrian Attributes Ontology
by: Ly, Ngoc Q., et al.
Published: (2025)
by: Ly, Ngoc Q., et al.
Published: (2025)
Leveraging Model Soups to Classify Intangible Cultural Heritage Images from the Mekong Delta
by: Tran, Quoc-Khang, et al.
Published: (2026)
by: Tran, Quoc-Khang, et al.
Published: (2026)
Adversarial Prompt Distillation for Vision-Language Models
by: Luo, Lin, et al.
Published: (2024)
by: Luo, Lin, et al.
Published: (2024)
Text-to-CT Generation via 3D Latent Diffusion Model with Contrastive Vision-Language Pretraining
by: Molino, Daniele, et al.
Published: (2025)
by: Molino, Daniele, et al.
Published: (2025)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
by: Liu, Yufang, et al.
Published: (2024)
by: Liu, Yufang, et al.
Published: (2024)
Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks
by: Song, Lingran, et al.
Published: (2025)
by: Song, Lingran, et al.
Published: (2025)
Multimodal Distribution Matching for Vision-Language Dataset Distillation
by: Jeong, Jongoh, et al.
Published: (2026)
by: Jeong, Jongoh, et al.
Published: (2026)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
by: Song, Chull Hwan, et al.
Published: (2024)
by: Song, Chull Hwan, et al.
Published: (2024)
MAE-Based Self-Supervised Pretraining for Data-Efficient Medical Image Segmentation Using nnFormer
by: Sureddi, R. M. Krishna, et al.
Published: (2026)
by: Sureddi, R. M. Krishna, et al.
Published: (2026)
Accuracy-Robustness Trade Off via Spiking Neural Network Gradient Sparsity Trail
by: Nhan, Luu Trong, et al.
Published: (2025)
by: Nhan, Luu Trong, et al.
Published: (2025)
Unifying Global and Local Scene Entities Modelling for Precise Action Spotting
by: Tran, Kim Hoang, et al.
Published: (2024)
by: Tran, Kim Hoang, et al.
Published: (2024)
Any-to-Any Learning in Computational Pathology via Triplet Multimodal Pretraining
by: Sun, Qichen, et al.
Published: (2025)
by: Sun, Qichen, et al.
Published: (2025)
Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular Diversity
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
ConPro: Learning Severity Representation for Medical Images using Contrastive Learning and Preference Optimization
by: Nguyen, Hong, et al.
Published: (2024)
by: Nguyen, Hong, et al.
Published: (2024)
Similar Items
-
Recent Advances in Medical Image Classification
by: Dao, Loan, et al.
Published: (2025) -
A Comprehensive Study on Medical Image Segmentation using Deep Neural Networks
by: Dao, Loan, et al.
Published: (2025) -
Enhancing Feature Diversity Boosts Channel-Adaptive Vision Transformers
by: Pham, Chau, et al.
Published: (2024) -
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
by: Tran, Quoc-Khang, et al.
Published: (2026) -
Ontology-based knowledge representation for bone disease diagnosis: a foundation for safe and sustainable medical artificial intelligence systems
by: Dao, Loan, et al.
Published: (2025)