MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Bowen, Ni, Minheng, Huang, Zitong, Yang, Guanglei, Zuo, Wangmeng, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MR-GDINO: Efficient Open-World Continual Object Detection
by: Dong, Bowen, et al.
Published: (2024)
by: Dong, Bowen, et al.
Published: (2024)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
ConSept: Continual Semantic Segmentation via Adapter-based Vision Transformer
by: Dong, Bowen, et al.
Published: (2024)
by: Dong, Bowen, et al.
Published: (2024)
LPT++: Efficient Training on Mixture of Long-tailed Experts
by: Dong, Bowen, et al.
Published: (2024)
by: Dong, Bowen, et al.
Published: (2024)
Responsible Visual Editing
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
by: Ni, Minheng, et al.
Published: (2025)
by: Ni, Minheng, et al.
Published: (2025)
Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
by: Zheng, Mi, et al.
Published: (2025)
by: Zheng, Mi, et al.
Published: (2025)
MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations
by: Dongre, Vardhan, et al.
Published: (2025)
by: Dongre, Vardhan, et al.
Published: (2025)
CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning
by: Yao, Zhenquan, et al.
Published: (2026)
by: Yao, Zhenquan, et al.
Published: (2026)
Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024)
by: Xu, Binqian, et al.
Published: (2024)
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
by: Ni, Minheng, et al.
Published: (2026)
by: Ni, Minheng, et al.
Published: (2026)
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
by: Ross, Candace, et al.
Published: (2025)
by: Ross, Candace, et al.
Published: (2025)
MetricDepth: Enhancing Monocular Depth Estimation with Deep Metric Learning
by: Liu, Chunpu, et al.
Published: (2024)
by: Liu, Chunpu, et al.
Published: (2024)
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
by: Sun, Linzhuang, et al.
Published: (2025)
by: Sun, Linzhuang, et al.
Published: (2025)
Don't Let Your Robot be Harmful: Responsible Robotic Manipulation via Safety-as-Policy
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
by: Zhang, Yabo, et al.
Published: (2025)
by: Zhang, Yabo, et al.
Published: (2025)
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
by: Zhang, Gengwei, et al.
Published: (2026)
by: Zhang, Gengwei, et al.
Published: (2026)
IMWA: Iterative Model Weight Averaging Benefits Class-Imbalanced Learning Tasks
by: Huang, Zitong, et al.
Published: (2024)
by: Huang, Zitong, et al.
Published: (2024)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning
by: Yerramilli, Sahiti, et al.
Published: (2025)
by: Yerramilli, Sahiti, et al.
Published: (2025)
MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
Improving Transferability of Adversarial Examples via Bayesian Attacks
by: Li, Qizhang, et al.
Published: (2023)
by: Li, Qizhang, et al.
Published: (2023)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
by: Li, Yi, et al.
Published: (2026)
by: Li, Yi, et al.
Published: (2026)
Latent Code Augmentation Based on Stable Diffusion for Data-free Substitute Attacks
by: Shao, Mingwen, et al.
Published: (2023)
by: Shao, Mingwen, et al.
Published: (2023)
FILP-3D: Enhancing 3D Few-shot Class-incremental Learning with Pre-trained Vision-Language Models
by: Xu, Wan, et al.
Published: (2023)
by: Xu, Wan, et al.
Published: (2023)
Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models
by: Dong, Yuhao, et al.
Published: (2026)
by: Dong, Yuhao, et al.
Published: (2026)
Multi-Modality Driven LoRA for Adverse Condition Depth Estimation
by: Yang, Guanglei, et al.
Published: (2024)
by: Yang, Guanglei, et al.
Published: (2024)
Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective
by: Li, Yuanze, et al.
Published: (2024)
by: Li, Yuanze, et al.
Published: (2024)
LLM as a Complementary Optimizer to Gradient Descent: A Case Study in Prompt Tuning
by: Guo, Zixian, et al.
Published: (2024)
by: Guo, Zixian, et al.
Published: (2024)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
by: Wang, Zhecan, et al.
Published: (2024)
by: Wang, Zhecan, et al.
Published: (2024)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
by: Villegas, Danae Sánchez, et al.
Published: (2025)
by: Villegas, Danae Sánchez, et al.
Published: (2025)
Steering the Verifiability of Multimodal AI Hallucinations
by: Pang, Jianhong, et al.
Published: (2026)
by: Pang, Jianhong, et al.
Published: (2026)
Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition
by: Wang, Zeheng, et al.
Published: (2026)
by: Wang, Zeheng, et al.
Published: (2026)
Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision-Language Models
by: Khanal, Bidur, et al.
Published: (2025)
by: Khanal, Bidur, et al.
Published: (2025)
Chain-of-Sketch: Enabling Global Visual Reasoning
by: Lotfi, Aryo, et al.
Published: (2024)
by: Lotfi, Aryo, et al.
Published: (2024)
HalluRNN: Mitigating Hallucinations via Recurrent Cross-Layer Reasoning in Large Vision-Language Models
by: Yu, Le, et al.
Published: (2025)
by: Yu, Le, et al.
Published: (2025)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
by: Li, Chengzu, et al.
Published: (2025)
by: Li, Chengzu, et al.
Published: (2025)
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
by: Li, Lingxiao, et al.
Published: (2025)
by: Li, Lingxiao, et al.
Published: (2025)
AutoDirector: Online Auto-scheduling Agents for Multi-sensory Composition
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
Similar Items
-
MR-GDINO: Efficient Open-World Continual Object Detection
by: Dong, Bowen, et al.
Published: (2024) -
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
by: Ni, Minheng, et al.
Published: (2024) -
ConSept: Continual Semantic Segmentation via Adapter-based Vision Transformer
by: Dong, Bowen, et al.
Published: (2024) -
LPT++: Efficient Training on Mixture of Long-tailed Experts
by: Dong, Bowen, et al.
Published: (2024) -
Responsible Visual Editing
by: Ni, Minheng, et al.
Published: (2024)