PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yantao, Hui, Qiang, Yan, Chenyang, Cheng, Kanzhi, Zhao, Fang, Tan, Chao, Gao, Huanling, Zhang, Jianbing, Wang, Kai, Dai, Xinyu, Lian, Shiguo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911493827592192
author Li, Yantao
Hui, Qiang
Yan, Chenyang
Cheng, Kanzhi
Zhao, Fang
Tan, Chao
Gao, Huanling
Zhang, Jianbing
Wang, Kai
Dai, Xinyu
Lian, Shiguo
author_facet Li, Yantao
Hui, Qiang
Yan, Chenyang
Cheng, Kanzhi
Zhao, Fang
Tan, Chao
Gao, Huanling
Zhang, Jianbing
Wang, Kai
Dai, Xinyu
Lian, Shiguo
contents Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucinations--cases where models reach the right answer while misperceiving visual evidence. We address this process-level misalignment with PaLMR, a framework that aligns not only outcomes but also the reasoning process itself. PaLMR comprises two complementary components: a perception-aligned data layer that constructs process-aware reasoning data with structured pseudo-ground-truths and verifiable visual facts, and a process-aligned optimisation layer that constructs a hierarchical reward fusion scheme with a process-aware scoring function to encourage visually faithful chains-of-thought and improve training stability. Experiments on Qwen2.5-VL-7B show that our approach substantially reduces reasoning hallucinations and improves visual reasoning fidelity, achieving state-of-the-art results on HallusionBench while maintaining strong performance on MMMU, MathVista, and MathVerse. These findings indicate that PaLMR offers a principled and practical route to process-aligned multimodal reasoning, advancing the reliability and interpretability of MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06652
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment
Li, Yantao
Hui, Qiang
Yan, Chenyang
Cheng, Kanzhi
Zhao, Fang
Tan, Chao
Gao, Huanling
Zhang, Jianbing
Wang, Kai
Dai, Xinyu
Lian, Shiguo
Computer Vision and Pattern Recognition
Artificial Intelligence
Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucinations--cases where models reach the right answer while misperceiving visual evidence. We address this process-level misalignment with PaLMR, a framework that aligns not only outcomes but also the reasoning process itself. PaLMR comprises two complementary components: a perception-aligned data layer that constructs process-aware reasoning data with structured pseudo-ground-truths and verifiable visual facts, and a process-aligned optimisation layer that constructs a hierarchical reward fusion scheme with a process-aware scoring function to encourage visually faithful chains-of-thought and improve training stability. Experiments on Qwen2.5-VL-7B show that our approach substantially reduces reasoning hallucinations and improves visual reasoning fidelity, achieving state-of-the-art results on HallusionBench while maintaining strong performance on MMMU, MathVista, and MathVerse. These findings indicate that PaLMR offers a principled and practical route to process-aligned multimodal reasoning, advancing the reliability and interpretability of MLLMs.
title PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.06652