Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Ce, Wan, Zifu, Kan, Zhehan, Ma, Martin Q., Stepputtis, Simon, Ramanan, Deva, Salakhutdinov, Russ, Morency, Louis-Philippe, Sycara, Katia, Xie, Yaqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
by: Wan, Zifu, et al.
Published: (2025)
by: Wan, Zifu, et al.
Published: (2025)
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
by: Wan, Zifu, et al.
Published: (2025)
by: Wan, Zifu, et al.
Published: (2025)
Spectral-Aware Global Fusion for RGB-Thermal Semantic Segmentation
by: Zhang, Ce, et al.
Published: (2025)
by: Zhang, Ce, et al.
Published: (2025)
Enhancing Vision-Language Few-Shot Adaptation with Negative Learning
by: Zhang, Ce, et al.
Published: (2024)
by: Zhang, Ce, et al.
Published: (2024)
Dual Prototype Evolving for Test-Time Generalization of Vision-Language Models
by: Zhang, Ce, et al.
Published: (2024)
by: Zhang, Ce, et al.
Published: (2024)
HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph Generation
by: Zhang, Ce, et al.
Published: (2024)
by: Zhang, Ce, et al.
Published: (2024)
Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation
by: Wan, Zifu, et al.
Published: (2024)
by: Wan, Zifu, et al.
Published: (2024)
GL-NeRF: Gauss-Laguerre Quadrature Enables Training-Free NeRF Acceleration
by: Yong, Silong, et al.
Published: (2024)
by: Yong, Silong, et al.
Published: (2024)
OMG: Opacity Matters in Material Modeling with Gaussian Splatting
by: Yong, Silong, et al.
Published: (2025)
by: Yong, Silong, et al.
Published: (2025)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023)
by: Mo, Shentong, et al.
Published: (2023)
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory
by: Zhang, Ce, et al.
Published: (2026)
by: Zhang, Ce, et al.
Published: (2026)
CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs
by: Kan, Zhehan, et al.
Published: (2024)
by: Kan, Zhehan, et al.
Published: (2024)
HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation
by: Wang, Zirui, et al.
Published: (2024)
by: Wang, Zirui, et al.
Published: (2024)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
Jailbreaking Frontier Foundation Models Through Intention Deception
by: Wang, Xinhe, et al.
Published: (2026)
by: Wang, Xinhe, et al.
Published: (2026)
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
by: Yuan, Fan, et al.
Published: (2024)
by: Yuan, Fan, et al.
Published: (2024)
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
by: Yu, Haofei, et al.
Published: (2023)
by: Yu, Haofei, et al.
Published: (2023)
Predicting Long-horizon Futures by Conditioning on Geometry and Time
by: Khurana, Tarasha, et al.
Published: (2024)
by: Khurana, Tarasha, et al.
Published: (2024)
Act2See: Emergent Active Visual Perception for Video Reasoning
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
Theory of Mind for Multi-Agent Collaboration via Large Language Models
by: Li, Huao, et al.
Published: (2023)
by: Li, Huao, et al.
Published: (2023)
Symbolic Graph Inference for Compound Scene Understanding
by: Aryan, FNU, et al.
Published: (2024)
by: Aryan, FNU, et al.
Published: (2024)
Exploring Causes and Mitigation of Hallucinations in Large Vision Language Models
by: Sun, Yaqi, et al.
Published: (2025)
by: Sun, Yaqi, et al.
Published: (2025)
Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition
by: Lee, Dong Won, et al.
Published: (2025)
by: Lee, Dong Won, et al.
Published: (2025)
Revisiting Few-Shot Object Detection with Vision-Language Models
by: Madan, Anish, et al.
Published: (2023)
by: Madan, Anish, et al.
Published: (2023)
RefAV: Towards Planning-Centric Scenario Mining
by: Davidson, Cainan, et al.
Published: (2025)
by: Davidson, Cainan, et al.
Published: (2025)
HEMM: Holistic Evaluation of Multimodal Foundation Models
by: Liang, Paul Pu, et al.
Published: (2024)
by: Liang, Paul Pu, et al.
Published: (2024)
MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
Revisiting the Role of Language Priors in Vision-Language Models
by: Lin, Zhiqiu, et al.
Published: (2023)
by: Lin, Zhiqiu, et al.
Published: (2023)
Multi-Agent Transfer Learning via Temporal Contrastive Learning
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
SMORE: Simultaneous Map and Object REconstruction
by: Chodosh, Nathaniel, et al.
Published: (2024)
by: Chodosh, Nathaniel, et al.
Published: (2024)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
by: Wang, Xintong, et al.
Published: (2024)
by: Wang, Xintong, et al.
Published: (2024)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
by: Lee, Dong Won, et al.
Published: (2024)
by: Lee, Dong Won, et al.
Published: (2024)
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
by: Mathur, Leena, et al.
Published: (2024)
by: Mathur, Leena, et al.
Published: (2024)
Efficient Contrastive Decoding with Probabilistic Hallucination Detection - Mitigating Hallucinations in Large Vision Language Models -
by: Fieback, Laura, et al.
Published: (2025)
by: Fieback, Laura, et al.
Published: (2025)
Using Diffusion Priors for Video Amodal Segmentation
by: Chen, Kaihua, et al.
Published: (2024)
by: Chen, Kaihua, et al.
Published: (2024)
Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
by: Chen, Kaihua, et al.
Published: (2025)
by: Chen, Kaihua, et al.
Published: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
by: Min, Kyungmin, et al.
Published: (2024)
by: Min, Kyungmin, et al.
Published: (2024)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
by: Lee, Yi-Lun, et al.
Published: (2024)
by: Lee, Yi-Lun, et al.
Published: (2024)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
Similar Items
-
ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
by: Wan, Zifu, et al.
Published: (2025) -
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
by: Wan, Zifu, et al.
Published: (2025) -
Spectral-Aware Global Fusion for RGB-Thermal Semantic Segmentation
by: Zhang, Ce, et al.
Published: (2025) -
Enhancing Vision-Language Few-Shot Adaptation with Negative Learning
by: Zhang, Ce, et al.
Published: (2024) -
Dual Prototype Evolving for Test-Time Generalization of Vision-Language Models
by: Zhang, Ce, et al.
Published: (2024)