Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Gengwei, Peng, Jie, Tan, Zhen, Qiu, Mufan, Mahjoub, Hossein Nourkhiz, Tadiparthi, Vaishnav, Lee, Kwonjoon, Zhang, Yanyong, Chen, Tianlong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Robust Reasoning through Guided Adversarial Self-Play
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction
by: Shen, Xu, et al.
Published: (2025)
by: Shen, Xu, et al.
Published: (2025)
Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication
by: Li, Huao, et al.
Published: (2024)
by: Li, Huao, et al.
Published: (2024)
Symbolic Graph Inference for Compound Scene Understanding
by: Aryan, FNU, et al.
Published: (2024)
by: Aryan, FNU, et al.
Published: (2024)
In Search of a Lost Metric: Human Empowerment as a Pillar of Socially Conscious Navigation
by: Baddam, Vasanth Reddy, et al.
Published: (2025)
by: Baddam, Vasanth Reddy, et al.
Published: (2025)
SSR: A Generic Framework for Text-Aided Map Compression for Localization
by: Omama, Mohammad, et al.
Published: (2026)
by: Omama, Mohammad, et al.
Published: (2026)
Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert Parallelism Design
by: Zhang, Mohan, et al.
Published: (2025)
by: Zhang, Mohan, et al.
Published: (2025)
GEM: 3D Gaussian Splatting for Efficient and Accurate Cryo-EM Reconstruction
by: Qu, Huaizhi, et al.
Published: (2025)
by: Qu, Huaizhi, et al.
Published: (2025)
MR-LDM -- The Merge-Reactive Longitudinal Decision Model: Game Theoretic Human Decision Modeling for Interactive Sim Agents
by: Holley, Dustin, et al.
Published: (2025)
by: Holley, Dustin, et al.
Published: (2025)
R3DM: Enabling Role Discovery and Diversity Through Dynamics Models in Multi-agent Reinforcement Learning
by: Goel, Harsh, et al.
Published: (2025)
by: Goel, Harsh, et al.
Published: (2025)
SMART-Merge Planner: A Safe Merging and Real-Time Motion Planner for Autonomous Highway On-Ramp Merging
by: Mohammadnejad, Toktam, et al.
Published: (2025)
by: Mohammadnejad, Toktam, et al.
Published: (2025)
Symbiotic Cooperation for Web Agents: Harnessing Complementary Strengths of Large and Small LLMs
by: Zhang, Ruichen, et al.
Published: (2025)
by: Zhang, Ruichen, et al.
Published: (2025)
Geometry- and Relation-Aware Diffusion for EEG Super-Resolution
by: Yao, Laura, et al.
Published: (2026)
by: Yao, Laura, et al.
Published: (2026)
Task-Aware Resolution Optimization for Visual Large Language Models
by: Luo, Weiqing, et al.
Published: (2025)
by: Luo, Weiqing, et al.
Published: (2025)
Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion
by: Knaup, Jacob, et al.
Published: (2025)
by: Knaup, Jacob, et al.
Published: (2025)
Vulnerability-Aware Robust Multimodal Adversarial Training
by: Zhang, Junrui, et al.
Published: (2025)
by: Zhang, Junrui, et al.
Published: (2025)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
by: Zhou, Kaiwen, et al.
Published: (2023)
by: Zhou, Kaiwen, et al.
Published: (2023)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
by: Li, Pingzhi, et al.
Published: (2024)
by: Li, Pingzhi, et al.
Published: (2024)
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
by: Ma, Ji, et al.
Published: (2026)
by: Ma, Ji, et al.
Published: (2026)
Modeling the Lane-Change Reactions to Merging Vehicles for Highway On-Ramp Simulations
by: Holley, Dustin, et al.
Published: (2024)
by: Holley, Dustin, et al.
Published: (2024)
SPATIA: Multimodal Generation and Prediction of Spatial Cell Phenotypes
by: Kong, Zhenglun, et al.
Published: (2025)
by: Kong, Zhenglun, et al.
Published: (2025)
OR-R1: Automating Modeling and Solving of Operations Research Optimization Problem via Test-Time Reinforcement Learning
by: Ding, Zezhen, et al.
Published: (2025)
by: Ding, Zezhen, et al.
Published: (2025)
Tuning-Free Accountable Intervention for LLM Deployment -- A Metacognitive Approach
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
GRNFormer: A Biologically-Guided Framework for Integrating Gene Regulatory Networks into RNA Foundation Models
by: Qiu, Mufan, et al.
Published: (2025)
by: Qiu, Mufan, et al.
Published: (2025)
Proactive aggression or passive resistance: A face perspective on why and how illegitimate tasks elicit various counterproductive work behaviours in employees
by: Fubin Jiang, et al.
Published: (2025)
by: Fubin Jiang, et al.
Published: (2025)
Can Hallucination Correction Improve Video-Language Alignment?
by: Zhao, Lingjun, et al.
Published: (2025)
by: Zhao, Lingjun, et al.
Published: (2025)
Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer
by: Qu, Huaizhi, et al.
Published: (2025)
by: Qu, Huaizhi, et al.
Published: (2025)
Chreode: A Cell World Model for One-Step Temporal Dynamics and Perturbation Prediction
by: Qiu, Mufan, et al.
Published: (2026)
by: Qiu, Mufan, et al.
Published: (2026)
Navigating Noisy Feedback: Enhancing Reinforcement Learning with Error-Prone Language Models
by: Lin, Muhan, et al.
Published: (2024)
by: Lin, Muhan, et al.
Published: (2024)
Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Harnessing Your DRAM and SSD for Sustainable and Accessible LLM Inference with Mixed-Precision and Multi-level Caching
by: Peng, Jie, et al.
Published: (2024)
by: Peng, Jie, et al.
Published: (2024)
Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
ATLS: Automated Trailer Loading for Surface Vessels
by: Abughaida, Amer, et al.
Published: (2024)
by: Abughaida, Amer, et al.
Published: (2024)
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
by: Tan, Zelin, et al.
Published: (2025)
by: Tan, Zelin, et al.
Published: (2025)
Task-aware Distributed Source Coding under Dynamic Bandwidth
by: Li, Po-han, et al.
Published: (2023)
by: Li, Po-han, et al.
Published: (2023)
ViTGAN: Training GANs with Vision Transformers
by: Lee, Kwonjoon, et al.
Published: (2021)
by: Lee, Kwonjoon, et al.
Published: (2021)
Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning
by: Lin, Muhan, et al.
Published: (2025)
by: Lin, Muhan, et al.
Published: (2025)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
TRUST: A Decentralized Framework for Auditing Large Language Model Reasoning
by: Huang, Morris Yu-Chao, et al.
Published: (2025)
by: Huang, Morris Yu-Chao, et al.
Published: (2025)
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
by: Dong, Bowen, et al.
Published: (2025)
by: Dong, Bowen, et al.
Published: (2025)
Similar Items
-
Learning Robust Reasoning through Guided Adversarial Self-Play
by: Li, Shuozhe, et al.
Published: (2026) -
Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction
by: Shen, Xu, et al.
Published: (2025) -
Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication
by: Li, Huao, et al.
Published: (2024) -
Symbolic Graph Inference for Compound Scene Understanding
by: Aryan, FNU, et al.
Published: (2024) -
In Search of a Lost Metric: Human Empowerment as a Pillar of Socially Conscious Navigation
by: Baddam, Vasanth Reddy, et al.
Published: (2025)