Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Pinxue, Wu, Chongruo, Zhou, Xinyu, Hong, Lingyi, Chen, Zhaoyu, Li, Jinglun, Jiang, Kaixun, Cheung, Sen-ching Samson, Zhang, Wei, Zhang, Wenqiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning
by: Li, Jinglun, et al.
Published: (2024)
by: Li, Jinglun, et al.
Published: (2024)
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
by: Hong, Lingyi, et al.
Published: (2026)
by: Hong, Lingyi, et al.
Published: (2026)
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
by: Fu, Jiyuan, et al.
Published: (2025)
by: Fu, Jiyuan, et al.
Published: (2025)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
Improving Adversarial Transferability with Neighbourhood Gradient Information
by: Guo, Haijing, et al.
Published: (2024)
by: Guo, Haijing, et al.
Published: (2024)
OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning
by: Hong, Lingyi, et al.
Published: (2024)
by: Hong, Lingyi, et al.
Published: (2024)
VideoPure: Diffusion-based Adversarial Purification for Video Recognition
by: Jiang, Kaixun, et al.
Published: (2025)
by: Jiang, Kaixun, et al.
Published: (2025)
Reading Relevant Feature from Global Representation Memory for Visual Object Tracking
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Boosting the Transferability of Adversarial Attacks with Global Momentum Initialization
by: Wang, Jiafeng, et al.
Published: (2022)
by: Wang, Jiafeng, et al.
Published: (2022)
General Compression Framework for Efficient Transformer Object Tracking
by: Hong, Lingyi, et al.
Published: (2024)
by: Hong, Lingyi, et al.
Published: (2024)
ClickVOS: Click Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
Boundary-Centric Active Learning for Temporal Action Segmentation
by: Helvaci, Halil Ismail, et al.
Published: (2026)
by: Helvaci, Halil Ismail, et al.
Published: (2026)
Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment
by: Jiang, Kaixun, et al.
Published: (2025)
by: Jiang, Kaixun, et al.
Published: (2025)
Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution Detection
by: Li, Jinglun, et al.
Published: (2024)
by: Li, Jinglun, et al.
Published: (2024)
Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection
by: Li, Jinglun, et al.
Published: (2025)
by: Li, Jinglun, et al.
Published: (2025)
LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation
by: Hong, Lingyi, et al.
Published: (2024)
by: Hong, Lingyi, et al.
Published: (2024)
VideoSAM: Open-World Video Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
OneVOS: Unifying Video Object Segmentation with All-in-One Transformer Framework
by: Li, Wanyun, et al.
Published: (2024)
by: Li, Wanyun, et al.
Published: (2024)
MMTA: Multi Membership Temporal Attention for Fine-Grained Stroke Rehabilitation Assessment
by: Helvaci, Halil Ismail, et al.
Published: (2026)
by: Helvaci, Halil Ismail, et al.
Published: (2026)
Exploring the Adversarial Robustness of Face Forgery Detection with Decision-based Black-box Attacks
by: Chen, Zhaoyu, et al.
Published: (2023)
by: Chen, Zhaoyu, et al.
Published: (2023)
VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking
by: Fu, Jiyuan, et al.
Published: (2026)
by: Fu, Jiyuan, et al.
Published: (2026)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
by: He, Zhentao, et al.
Published: (2025)
by: He, Zhentao, et al.
Published: (2025)
HRTR: A Single-stage Transformer for Fine-grained Sub-second Action Segmentation in Stroke Rehabilitation
by: Helvaci, Halil Ismail, et al.
Published: (2025)
by: Helvaci, Halil Ismail, et al.
Published: (2025)
Seeing is Believing? Enhancing Vision-Language Navigation using Visual Perturbations
by: Zhang, Xuesong, et al.
Published: (2024)
by: Zhang, Xuesong, et al.
Published: (2024)
Localizing Moments of Actions in Untrimmed Videos of Infants with Autism Spectrum Disorder
by: Helvaci, Halil Ismail, et al.
Published: (2024)
by: Helvaci, Halil Ismail, et al.
Published: (2024)
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
by: Ling, Chen, et al.
Published: (2026)
by: Ling, Chen, et al.
Published: (2026)
Visual Room 2.0: Seeing is Not Understanding for MLLMs
by: Li, Haokun, et al.
Published: (2025)
by: Li, Haokun, et al.
Published: (2025)
Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback, Diversity-Aware Sampling, and Conflict Regularization
by: Pan, Miao, et al.
Published: (2026)
by: Pan, Miao, et al.
Published: (2026)
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
by: Narnaware, Vishal, et al.
Published: (2026)
by: Narnaware, Vishal, et al.
Published: (2026)
PG-Attack: A Precision-Guided Adversarial Attack Framework Against Vision Foundation Models for Autonomous Driving
by: Fu, Jiyuan, et al.
Published: (2024)
by: Fu, Jiyuan, et al.
Published: (2024)
OpenVIS: Open-vocabulary Video Instance Segmentation
by: Guo, Pinxue, et al.
Published: (2023)
by: Guo, Pinxue, et al.
Published: (2023)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
by: Deng, Ailin, et al.
Published: (2024)
by: Deng, Ailin, et al.
Published: (2024)
Seeing is not Believing: An Identity Hider for Human Vision Privacy Protection
by: Wang, Tao, et al.
Published: (2023)
by: Wang, Tao, et al.
Published: (2023)
Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey
by: Zhang, Xiantao
Published: (2025)
by: Zhang, Xiantao
Published: (2025)
RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations
by: He, Xingqi, et al.
Published: (2025)
by: He, Xingqi, et al.
Published: (2025)
See the Unseen: Better Context-Consistent Knowledge-Editing by Noises
by: Huang, Youcheng, et al.
Published: (2024)
by: Huang, Youcheng, et al.
Published: (2024)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
by: Tang, Feilong, et al.
Published: (2025)
by: Tang, Feilong, et al.
Published: (2025)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
by: Ji, Yikun, et al.
Published: (2025)
by: Ji, Yikun, et al.
Published: (2025)
Similar Items
-
TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning
by: Li, Jinglun, et al.
Published: (2024) -
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking
by: Zhou, Xinyu, et al.
Published: (2025) -
Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
by: Hong, Lingyi, et al.
Published: (2026) -
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
by: Fu, Jiyuan, et al.
Published: (2025) -
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)