Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
Fuente:
arXiv
Saved in:
| Main Authors: | Moon, Jihyun, Hong, Charmgil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning-Augmented Representations for Multimodal Retrieval
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Open Multimodal Retrieval-Augmented Factual Image Generation
by: Tian, Yang, et al.
Published: (2025)
by: Tian, Yang, et al.
Published: (2025)
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
by: Mei, Jingbiao, et al.
Published: (2025)
by: Mei, Jingbiao, et al.
Published: (2025)
Rethinking VLMs and LLMs for Image Classification
by: Cooper, Avi, et al.
Published: (2024)
by: Cooper, Avi, et al.
Published: (2024)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)
by: Li, Shuo, et al.
Published: (2024)
RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models
by: Hao, Haoran, et al.
Published: (2024)
by: Hao, Haoran, et al.
Published: (2024)
CARES: Context-Aware Resolution Selector for VLMs
by: Kimhi, Moshe, et al.
Published: (2025)
by: Kimhi, Moshe, et al.
Published: (2025)
DASH: Detection and Assessment of Systematic Hallucinations of VLMs
by: Augustin, Maximilian, et al.
Published: (2025)
by: Augustin, Maximilian, et al.
Published: (2025)
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
by: Liu, Ziyu, et al.
Published: (2024)
by: Liu, Ziyu, et al.
Published: (2024)
Hidden in plain sight: VLMs overlook their visual representations
by: Fu, Stephanie, et al.
Published: (2025)
by: Fu, Stephanie, et al.
Published: (2025)
Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight
by: Ding, Xi, et al.
Published: (2024)
by: Ding, Xi, et al.
Published: (2024)
DEAL: Disentangle and Localize Concept-level Explanations for VLMs
by: Li, Tang, et al.
Published: (2024)
by: Li, Tang, et al.
Published: (2024)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
by: Ming, Yifei, et al.
Published: (2024)
by: Ming, Yifei, et al.
Published: (2024)
Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
by: Faraz, Ali, et al.
Published: (2025)
by: Faraz, Ali, et al.
Published: (2025)
Do We Need Large VLMs for Spotting Soccer Actions?
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
by: Chen, Weixin, et al.
Published: (2026)
by: Chen, Weixin, et al.
Published: (2026)
MMM: Quantum-Chemical Molecular Representation Learning for Combinatorial Drug Recommendation
by: Kwon, Chongmyung, et al.
Published: (2025)
by: Kwon, Chongmyung, et al.
Published: (2025)
Learning Superpixel Ensemble and Hierarchy Graphs for Melanoma Detection
by: Elwer, Asmaa M., et al.
Published: (2026)
by: Elwer, Asmaa M., et al.
Published: (2026)
Few-Shot Recognition via Stage-Wise Retrieval-Augmented Finetuning
by: Liu, Tian, et al.
Published: (2024)
by: Liu, Tian, et al.
Published: (2024)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025)
by: Izadi, Amirmohammad, et al.
Published: (2025)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024)
by: Li, Kevin Y., et al.
Published: (2024)
COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
by: Xia, Canming, et al.
Published: (2026)
by: Xia, Canming, et al.
Published: (2026)
Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks
by: Matsuishi, Koki, et al.
Published: (2025)
by: Matsuishi, Koki, et al.
Published: (2025)
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
by: Kumar, Sunil, et al.
Published: (2025)
by: Kumar, Sunil, et al.
Published: (2025)
Right this way: Can VLMs Guide Us to See More to Answer Questions?
by: Liu, Li, et al.
Published: (2024)
by: Liu, Li, et al.
Published: (2024)
Adapting Segment Anything Model to Melanoma Segmentation in Microscopy Slide Images
by: Liu, Qingyuan, et al.
Published: (2024)
by: Liu, Qingyuan, et al.
Published: (2024)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation
by: Luo, Weiqing, et al.
Published: (2026)
by: Luo, Weiqing, et al.
Published: (2026)
Weakly Supervised Pretraining and Multi-Annotator Supervised Finetuning for Facial Wrinkle Detection
by: Moon, Ik Jun, et al.
Published: (2024)
by: Moon, Ik Jun, et al.
Published: (2024)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
by: Pantazopoulos, Georgios, et al.
Published: (2024)
by: Pantazopoulos, Georgios, et al.
Published: (2024)
Can World Models Benefit VLMs for World Dynamics?
by: Zhang, Kevin, et al.
Published: (2025)
by: Zhang, Kevin, et al.
Published: (2025)
Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy
by: Bouguerra, Aymen, et al.
Published: (2025)
by: Bouguerra, Aymen, et al.
Published: (2025)
Frugal Federated Learning for Violence Detection: A Comparison of LoRA-Tuned VLMs and Personalized CNNs
by: Thuau, Sébastien, et al.
Published: (2025)
by: Thuau, Sébastien, et al.
Published: (2025)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
by: Pan, Zhiyu, et al.
Published: (2026)
by: Pan, Zhiyu, et al.
Published: (2026)
A Multimodal RAG Framework for Housing Damage Assessment: Collaborative Optimization of Image Encoding and Policy Vector Retrieval
by: Miao, Jiayi, et al.
Published: (2025)
by: Miao, Jiayi, et al.
Published: (2025)
Sample Selection via Contrastive Fragmentation for Noisy Label Regression
by: Kim, Chris Dongjoo, et al.
Published: (2025)
by: Kim, Chris Dongjoo, et al.
Published: (2025)
Similar Items
-
Reasoning-Augmented Representations for Multimodal Retrieval
by: Zhang, Jianrui, et al.
Published: (2026) -
Open Multimodal Retrieval-Augmented Factual Image Generation
by: Tian, Yang, et al.
Published: (2025) -
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
by: Mei, Jingbiao, et al.
Published: (2025) -
Rethinking VLMs and LLMs for Image Classification
by: Cooper, Avi, et al.
Published: (2024) -
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)