Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wu, Yuanchen, Yan, Ke, Ding, Shouhong, Zhou, Ziyin, Li, Xiaoqiang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911159282565120
author Wu, Yuanchen
Yan, Ke
Ding, Shouhong
Zhou, Ziyin
Li, Xiaoqiang
author_facet Wu, Yuanchen
Yan, Ke
Ding, Shouhong
Zhou, Ziyin
Li, Xiaoqiang
contents Large Vision-Language Models (LVLMs) have manifested strong visual question answering capability. However, they still struggle with aligning the rationale and the generated answer, leading to inconsistent reasoning and incorrect responses. To this end, this paper introduces the Self-Rationale Calibration (SRC) framework to iteratively calibrate the alignment between rationales and answers. SRC begins by employing a lightweight "rationale fine-tuning" approach, which modifies the model's response format to require a rationale before deriving an answer without explicit prompts. Next, SRC searches for a diverse set of candidate responses from the fine-tuned LVLMs for each sample, followed by a proposed pairwise scoring strategy using a tailored scoring model, R-Scorer, to evaluate both rationale quality and factual consistency of candidates. Based on a confidence-weighted preference curation process, SRC decouples the alignment calibration into a preference fine-tuning manner, leading to significant improvements of LVLMs in perception, reasoning, and generalization across multiple benchmarks. Our results emphasize the rationale-oriented alignment in exploring the potential of LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13919
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
Wu, Yuanchen
Yan, Ke
Ding, Shouhong
Zhou, Ziyin
Li, Xiaoqiang
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) have manifested strong visual question answering capability. However, they still struggle with aligning the rationale and the generated answer, leading to inconsistent reasoning and incorrect responses. To this end, this paper introduces the Self-Rationale Calibration (SRC) framework to iteratively calibrate the alignment between rationales and answers. SRC begins by employing a lightweight "rationale fine-tuning" approach, which modifies the model's response format to require a rationale before deriving an answer without explicit prompts. Next, SRC searches for a diverse set of candidate responses from the fine-tuned LVLMs for each sample, followed by a proposed pairwise scoring strategy using a tailored scoring model, R-Scorer, to evaluate both rationale quality and factual consistency of candidates. Based on a confidence-weighted preference curation process, SRC decouples the alignment calibration into a preference fine-tuning manner, leading to significant improvements of LVLMs in perception, reasoning, and generalization across multiple benchmarks. Our results emphasize the rationale-oriented alignment in exploring the potential of LVLMs.
title Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.13919