Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback, Diversity-Aware Sampling, and Conflict Regularization
Fuente:
arXiv
Salvato in:
| Autori principali: | Pan, Miao, Gan, Wangjie, Chen, Jintao, Zhang, Wenqi, Sun, Bing, Yin, Jianwei, Zhang, Xuhong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
di: Gan, Wangjie, et al.
Pubblicazione: (2026)
di: Gan, Wangjie, et al.
Pubblicazione: (2026)
SCHK-HTC: Sibling Contrastive Learning with Hierarchical Knowledge-Aware Prompt Tuning for Hierarchical Text Classification
di: Xiong, Ke, et al.
Pubblicazione: (2026)
di: Xiong, Ke, et al.
Pubblicazione: (2026)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
di: Guo, Pinxue, et al.
Pubblicazione: (2025)
di: Guo, Pinxue, et al.
Pubblicazione: (2025)
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
di: Liu, Yanming, et al.
Pubblicazione: (2026)
di: Liu, Yanming, et al.
Pubblicazione: (2026)
See or Guess: Counterfactually Regularized Image Captioning
di: Cao, Qian, et al.
Pubblicazione: (2024)
di: Cao, Qian, et al.
Pubblicazione: (2024)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
di: Tang, Feilong, et al.
Pubblicazione: (2025)
di: Tang, Feilong, et al.
Pubblicazione: (2025)
Vision Language Models See What You Want but not What You See
di: Gao, Qingying, et al.
Pubblicazione: (2024)
di: Gao, Qingying, et al.
Pubblicazione: (2024)
"Can You See Me Think?" Grounding LLM Feedback in Keystrokes and Revision Patterns
di: Zafar, Samra, et al.
Pubblicazione: (2025)
di: Zafar, Samra, et al.
Pubblicazione: (2025)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
di: Jiang, Yankai, et al.
Pubblicazione: (2026)
di: Jiang, Yankai, et al.
Pubblicazione: (2026)
What You See Is What You Get: Attention-based Self-guided Automatic Unit Test Generation
di: Yin, Xin, et al.
Pubblicazione: (2024)
di: Yin, Xin, et al.
Pubblicazione: (2024)
RA-ISF: Learning to Answer and Understand from Retrieval Augmentation via Iterative Self-Feedback
di: Liu, Yanming, et al.
Pubblicazione: (2024)
di: Liu, Yanming, et al.
Pubblicazione: (2024)
Sample from What You See: Visuomotor Policy Learning via Diffusion Bridge with Observation-Embedded Stochastic Differential Equation
di: Liu, Zhaoyang, et al.
Pubblicazione: (2025)
di: Liu, Zhaoyang, et al.
Pubblicazione: (2025)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
di: Li, Qiaoru, et al.
Pubblicazione: (2026)
di: Li, Qiaoru, et al.
Pubblicazione: (2026)
What You See is What You Ask: Evaluating Audio Descriptions
di: Kala, Divy, et al.
Pubblicazione: (2025)
di: Kala, Divy, et al.
Pubblicazione: (2025)
What You See is What You Classify: Black Box Attributions
di: Stalder, Steven, et al.
Pubblicazione: (2022)
di: Stalder, Steven, et al.
Pubblicazione: (2022)
Do MLLMs See What We See? Analyzing Visualization Literacy Barriers in AI Systems
di: Mengli, et al.
Pubblicazione: (2026)
di: Mengli, et al.
Pubblicazione: (2026)
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
di: Narnaware, Vishal, et al.
Pubblicazione: (2026)
di: Narnaware, Vishal, et al.
Pubblicazione: (2026)
Visual Room 2.0: Seeing is Not Understanding for MLLMs
di: Li, Haokun, et al.
Pubblicazione: (2025)
di: Li, Haokun, et al.
Pubblicazione: (2025)
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
di: Yin, Zhihan, et al.
Pubblicazione: (2026)
di: Yin, Zhihan, et al.
Pubblicazione: (2026)
Spatially Selective Imaging in Color: What You See is What You Want
di: John You En Chan, et al.
Pubblicazione: (2024)
di: John You En Chan, et al.
Pubblicazione: (2024)
Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions
di: Bao, Yuntai, et al.
Pubblicazione: (2026)
di: Bao, Yuntai, et al.
Pubblicazione: (2026)
SARE: Sample-wise Adaptive Reasoning for Training-free Fine-grained Visual Recognition
di: Yang, Jingxiao, et al.
Pubblicazione: (2026)
di: Yang, Jingxiao, et al.
Pubblicazione: (2026)
What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
di: Li, Haoxi, et al.
Pubblicazione: (2026)
di: Li, Haoxi, et al.
Pubblicazione: (2026)
Explore the Hallucination on Low-level Perception for MLLMs
di: Sun, Yinan, et al.
Pubblicazione: (2024)
di: Sun, Yinan, et al.
Pubblicazione: (2024)
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
di: Lin, Weifeng, et al.
Pubblicazione: (2024)
di: Lin, Weifeng, et al.
Pubblicazione: (2024)
Say What You Mean: Natural Language Access Control with Large Language Models for Internet of Things
di: Cheng, Ye, et al.
Pubblicazione: (2025)
di: Cheng, Ye, et al.
Pubblicazione: (2025)
Neural Diversity Regularizes Hallucinations in Language Models
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2025)
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2025)
Mitigating Image Captioning Hallucinations in Vision-Language Models
di: Zhao, Fei, et al.
Pubblicazione: (2025)
di: Zhao, Fei, et al.
Pubblicazione: (2025)
Image Embedding Sampling Method for Diverse Captioning
di: Waheed, Sania, et al.
Pubblicazione: (2025)
di: Waheed, Sania, et al.
Pubblicazione: (2025)
Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions
di: Bao, Yuntai, et al.
Pubblicazione: (2026)
di: Bao, Yuntai, et al.
Pubblicazione: (2026)
What You See is Not What You Get: Neural Partial Differential Equations and The Illusion of Learning
di: Mohan, Arvind, et al.
Pubblicazione: (2024)
di: Mohan, Arvind, et al.
Pubblicazione: (2024)
What You See Is Not Always What You Get: Evaluating GPT's Comprehension of Source Code
di: Wen, Jiawen, et al.
Pubblicazione: (2024)
di: Wen, Jiawen, et al.
Pubblicazione: (2024)
LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning
di: Chao, Lianying, et al.
Pubblicazione: (2026)
di: Chao, Lianying, et al.
Pubblicazione: (2026)
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
di: Dong, Zixuan, et al.
Pubblicazione: (2025)
di: Dong, Zixuan, et al.
Pubblicazione: (2025)
Aesthetic Image Captioning with Saliency Enhanced MLLMs
di: Tao, Yilin, et al.
Pubblicazione: (2025)
di: Tao, Yilin, et al.
Pubblicazione: (2025)
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
di: Zhang, Qizhe, et al.
Pubblicazione: (2025)
di: Zhang, Qizhe, et al.
Pubblicazione: (2025)
PragLocker: Protecting Agent Intellectual Property in Untrusted Deployments via Non-Portable Prompts
di: Li, Qinfeng, et al.
Pubblicazione: (2026)
di: Li, Qinfeng, et al.
Pubblicazione: (2026)
Trusting What You Cannot See: Auditable Fine-Tuning and Inference for Proprietary AI
di: Jin, Heng, et al.
Pubblicazione: (2026)
di: Jin, Heng, et al.
Pubblicazione: (2026)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs
di: Tasnim, Nazia, et al.
Pubblicazione: (2025)
di: Tasnim, Nazia, et al.
Pubblicazione: (2025)
VERIRL: Boosting the LLM-based Verilog Code Generation via Reinforcement Learning
di: Teng, Fu, et al.
Pubblicazione: (2025)
di: Teng, Fu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
di: Gan, Wangjie, et al.
Pubblicazione: (2026) -
SCHK-HTC: Sibling Contrastive Learning with Hierarchical Knowledge-Aware Prompt Tuning for Hierarchical Text Classification
di: Xiong, Ke, et al.
Pubblicazione: (2026) -
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
di: Guo, Pinxue, et al.
Pubblicazione: (2025) -
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
di: Liu, Yanming, et al.
Pubblicazione: (2026) -
See or Guess: Counterfactually Regularized Image Captioning
di: Cao, Qian, et al.
Pubblicazione: (2024)