HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Songtao, Zhang, Yan, Jin, Yeying, Tang, Zhihang, Wu, Yangyang, Feng, Yang, Wu, Jian, Liu, Zuozhu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913869769736192
author Jiang, Songtao
Zhang, Yan
Jin, Yeying
Tang, Zhihang
Wu, Yangyang
Feng, Yang
Wu, Jian
Liu, Zuozhu
author_facet Jiang, Songtao
Zhang, Yan
Jin, Yeying
Tang, Zhihang
Wu, Yangyang
Feng, Yang
Wu, Jian
Liu, Zuozhu
contents Medical Vision-Language Models (Med-VLMs) have achieved success across various tasks, yet most existing methods overlook the modality misalignment issue that can lead to untrustworthy responses in clinical settings. In this paper, we propose Hierarchical Self-Contrastive Rewarding (HSCR), a novel approach that addresses two critical challenges in Med-VLM alignment: 1) Cost-effective generation of high-quality preference data; 2) Capturing nuanced and context-aware preferences for improved alignment. HSCR first leverages the inherent capability of Med-VLMs to generate dispreferred responses with higher sampling probability. By analyzing output logit shifts after visual token dropout, we identify modality-coupled tokens that induce misalignment and derive an implicit alignment reward function. This function guides token replacement with hallucinated ones during decoding, producing high-quality dispreferred data. Furthermore, HSCR introduces a multi-level preference optimization strategy, which extends beyond traditional adjacent-level optimization by incorporating nuanced implicit preferences, leveraging relative quality in dispreferred data to capture subtle alignment cues for more precise and context-aware optimization. Extensive experiments across multiple medical tasks, including Med-VQA, medical image captioning and instruction following, demonstrate that HSCR not only enhances zero-shot performance but also significantly improves modality alignment and trustworthiness with just 2,000 training entries.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
Jiang, Songtao
Zhang, Yan
Jin, Yeying
Tang, Zhihang
Wu, Yangyang
Feng, Yang
Wu, Jian
Liu, Zuozhu
Computer Vision and Pattern Recognition
Computation and Language
Medical Vision-Language Models (Med-VLMs) have achieved success across various tasks, yet most existing methods overlook the modality misalignment issue that can lead to untrustworthy responses in clinical settings. In this paper, we propose Hierarchical Self-Contrastive Rewarding (HSCR), a novel approach that addresses two critical challenges in Med-VLM alignment: 1) Cost-effective generation of high-quality preference data; 2) Capturing nuanced and context-aware preferences for improved alignment. HSCR first leverages the inherent capability of Med-VLMs to generate dispreferred responses with higher sampling probability. By analyzing output logit shifts after visual token dropout, we identify modality-coupled tokens that induce misalignment and derive an implicit alignment reward function. This function guides token replacement with hallucinated ones during decoding, producing high-quality dispreferred data. Furthermore, HSCR introduces a multi-level preference optimization strategy, which extends beyond traditional adjacent-level optimization by incorporating nuanced implicit preferences, leveraging relative quality in dispreferred data to capture subtle alignment cues for more precise and context-aware optimization. Extensive experiments across multiple medical tasks, including Med-VQA, medical image captioning and instruction following, demonstrate that HSCR not only enhances zero-shot performance but also significantly improves modality alignment and trustworthiness with just 2,000 training entries.
title HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2506.00805