PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yantao, Hui, Qiang, Yan, Chenyang, Cheng, Kanzhi, Zhao, Fang, Tan, Chao, Gao, Huanling, Zhang, Jianbing, Wang, Kai, Dai, Xinyu, Lian, Shiguo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-Language Models Can Self-Improve Reasoning via Reflection
by: Cheng, Kanzhi, et al.
Published: (2024)
by: Cheng, Kanzhi, et al.
Published: (2024)
MediaClaw: Multimodal Intelligent-Agent Platform Technical Report
by: Zhao, Shaoan, et al.
Published: (2026)
by: Zhao, Shaoan, et al.
Published: (2026)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024)
by: Cheng, Kanzhi, et al.
Published: (2024)
Unsupervised Industrial Anomaly Detection via Pattern Generative and Contrastive Networks
by: Huang, Jianfeng, et al.
Published: (2022)
by: Huang, Jianfeng, et al.
Published: (2022)
Patch-wise Auto-Encoder for Visual Anomaly Detection
by: Cui, Yajie, et al.
Published: (2023)
by: Cui, Yajie, et al.
Published: (2023)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
by: Wu, Ruijia, et al.
Published: (2025)
by: Wu, Ruijia, et al.
Published: (2025)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
by: Tian, Changyuan, et al.
Published: (2026)
by: Tian, Changyuan, et al.
Published: (2026)
Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description
by: Pan, Mianzhi, et al.
Published: (2023)
by: Pan, Mianzhi, et al.
Published: (2023)
A Large Vision-Language Model based Environment Perception System for Visually Impaired People
by: Chen, Zezhou, et al.
Published: (2025)
by: Chen, Zezhou, et al.
Published: (2025)
LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
by: Gao, Huanlin, et al.
Published: (2025)
by: Gao, Huanlin, et al.
Published: (2025)
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
by: Shen, Xu, et al.
Published: (2025)
by: Shen, Xu, et al.
Published: (2025)
RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal Evaluation
by: Wang, Haofeng, et al.
Published: (2025)
by: Wang, Haofeng, et al.
Published: (2025)
A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
by: Yan, Shuo, et al.
Published: (2025)
by: Yan, Shuo, et al.
Published: (2025)
Transfer-LMR: Heavy-Tail Driving Behavior Recognition in Diverse Traffic Scenarios
by: Parikh, Chirag, et al.
Published: (2024)
by: Parikh, Chirag, et al.
Published: (2024)
Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
by: Imam, Mohamed Fazli, et al.
Published: (2025)
by: Imam, Mohamed Fazli, et al.
Published: (2025)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
by: Cheng, Kanzhi, et al.
Published: (2025)
by: Cheng, Kanzhi, et al.
Published: (2025)
Towards Faithful Multimodal Concept Bottleneck Models
by: Moreau, Pierre, et al.
Published: (2026)
by: Moreau, Pierre, et al.
Published: (2026)
The Devil is in the Few Shots: Iterative Visual Knowledge Completion for Few-shot Learning
by: Li, Yaohui, et al.
Published: (2024)
by: Li, Yaohui, et al.
Published: (2024)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
by: Wang, Weiyun, et al.
Published: (2025)
by: Wang, Weiyun, et al.
Published: (2025)
Towards Faithful Reasoning in Comics for Small MLLMs
by: Feng, Chengcheng, et al.
Published: (2026)
by: Feng, Chengcheng, et al.
Published: (2026)
Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
by: Li, Junxian, et al.
Published: (2025)
by: Li, Junxian, et al.
Published: (2025)
Data-Driven Deepfake Image Detection Method -- The 2024 Global Deepfake Image Detection Challenge
by: Zhu, Xiaoya, et al.
Published: (2025)
by: Zhu, Xiaoya, et al.
Published: (2025)
TP3M: Transformer-based Pseudo 3D Image Matching with Reference Image
by: Han, Liming, et al.
Published: (2024)
by: Han, Liming, et al.
Published: (2024)
Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
PROPA: Toward Process-level Optimization in Visual Reasoning via Reinforcement Learning
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
KAConvNet: Kolmogorov-Arnold Convolutional Networks for Vision Recognition
by: Liu, Zhaoxiang, et al.
Published: (2026)
by: Liu, Zhaoxiang, et al.
Published: (2026)
Piculet: Specialized Models-Guided Hallucination Decrease for MultiModal Large Language Models
by: Wang, Kohou, et al.
Published: (2024)
by: Wang, Kohou, et al.
Published: (2024)
Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning
by: Luo, Linhao, et al.
Published: (2023)
by: Luo, Linhao, et al.
Published: (2023)
MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching Inference
by: Gao, Huanlin, et al.
Published: (2026)
by: Gao, Huanlin, et al.
Published: (2026)
LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data
by: Yang, Cehao, et al.
Published: (2025)
by: Yang, Cehao, et al.
Published: (2025)
Towards Efficient Visual-Language Alignment of the Q-Former for Visual Reasoning Tasks
by: Kim, Sungkyung, et al.
Published: (2024)
by: Kim, Sungkyung, et al.
Published: (2024)
Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning
by: Lee, Naeun, et al.
Published: (2026)
by: Lee, Naeun, et al.
Published: (2026)
Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
by: Cousins, Cyrus, et al.
Published: (2025)
by: Cousins, Cyrus, et al.
Published: (2025)
Fuzzy Reasoning Chain (FRC): An Innovative Reasoning Framework from Fuzziness to Clarity
by: Chen, Ping, et al.
Published: (2025)
by: Chen, Ping, et al.
Published: (2025)
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
by: Yang, Tianyu, et al.
Published: (2026)
by: Yang, Tianyu, et al.
Published: (2026)
Algorithms for variational Monte Carlo calculations of fermion projected entangled pair states in the swap gates formulation and the detailed balance of tensor network sequential sampling
by: Wu, Yantao, et al.
Published: (2025)
by: Wu, Yantao, et al.
Published: (2025)
Diffusion Model with Representation Alignment for Protein Inverse Folding
by: Wang, Chenglin, et al.
Published: (2024)
by: Wang, Chenglin, et al.
Published: (2024)
Similar Items
-
Vision-Language Models Can Self-Improve Reasoning via Reflection
by: Cheng, Kanzhi, et al.
Published: (2024) -
MediaClaw: Multimodal Intelligent-Agent Platform Technical Report
by: Zhao, Shaoan, et al.
Published: (2026) -
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024) -
Unsupervised Industrial Anomaly Detection via Pattern Generative and Contrastive Networks
by: Huang, Jianfeng, et al.
Published: (2022) -
Patch-wise Auto-Encoder for Visual Anomaly Detection
by: Cui, Yajie, et al.
Published: (2023)