Multi-Modal Hallucination Control by Visual Information Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Favero, Alessandro, Zancato, Luca, Trager, Matthew, Choudhary, Siddharth, Perera, Pramuditha, Achille, Alessandro, Swaminathan, Ashwin, Soatto, Stefano |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
by: Trager, Matthew, et al.
Published: (2023)
by: Trager, Matthew, et al.
Published: (2023)
Descriminative-Generative Custom Tokens for Vision-Language Models
by: Perera, Pramuditha, et al.
Published: (2025)
by: Perera, Pramuditha, et al.
Published: (2025)
Compositional Structures in Neural Embedding and Interaction Decompositions
by: Trager, Matthew, et al.
Published: (2024)
by: Trager, Matthew, et al.
Published: (2024)
PICASO: Permutation-Invariant Context Composition with State Space Models
by: Liu, Tian Yu, et al.
Published: (2025)
by: Liu, Tian Yu, et al.
Published: (2025)
CPR: Retrieval Augmented Generation for Copyright Protection
by: Golatkar, Aditya, et al.
Published: (2024)
by: Golatkar, Aditya, et al.
Published: (2024)
NeRF-Insert: 3D Local Editing with Multimodal Control Signals
by: Sabat, Benet Oriol, et al.
Published: (2024)
by: Sabat, Benet Oriol, et al.
Published: (2024)
Training Data Protection with Compositional Diffusion Models
by: Golatkar, Aditya, et al.
Published: (2023)
by: Golatkar, Aditya, et al.
Published: (2023)
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
by: Zancato, Luca, et al.
Published: (2024)
by: Zancato, Luca, et al.
Published: (2024)
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
by: Choudhary, Sakshi, et al.
Published: (2026)
by: Choudhary, Sakshi, et al.
Published: (2026)
Maximally-Informative Retrieval for State Space Model Generation
by: Becker, Evan, et al.
Published: (2025)
by: Becker, Evan, et al.
Published: (2025)
Interpretable Measures of Conceptual Similarity by Complexity-Constrained Descriptive Auto-Encoding
by: Achille, Alessandro, et al.
Published: (2024)
by: Achille, Alessandro, et al.
Published: (2024)
Grounded Visual Factualization: Factual Anchor-Based Finetuning for Enhancing MLLM Factual Consistency
by: Morbiato, Filippo, et al.
Published: (2025)
by: Morbiato, Filippo, et al.
Published: (2025)
Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression
by: Peng, Liangzu, et al.
Published: (2025)
by: Peng, Liangzu, et al.
Published: (2025)
Expansion Span: Combining Fading Memory and Retrieval in Hybrid State Space Models
by: Nunez, Elvis, et al.
Published: (2024)
by: Nunez, Elvis, et al.
Published: (2024)
Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
by: Biggs, Benjamin, et al.
Published: (2024)
by: Biggs, Benjamin, et al.
Published: (2024)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
by: Ghazanfari, Sara, et al.
Published: (2024)
by: Ghazanfari, Sara, et al.
Published: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
by: Kaul, Prannay, et al.
Published: (2024)
by: Kaul, Prannay, et al.
Published: (2024)
Mixed-Query Transformer: A Unified Image Segmentation Architecture
by: Wang, Pei, et al.
Published: (2024)
by: Wang, Pei, et al.
Published: (2024)
e1: Learning Adaptive Control of Reasoning Effort
by: Kleinman, Michael, et al.
Published: (2025)
by: Kleinman, Michael, et al.
Published: (2025)
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
by: Li, Xiaolong, et al.
Published: (2024)
by: Li, Xiaolong, et al.
Published: (2024)
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
by: Rajabi, Navid, et al.
Published: (2023)
by: Rajabi, Navid, et al.
Published: (2023)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
by: Li, Zejun, et al.
Published: (2024)
by: Li, Zejun, et al.
Published: (2024)
Fast Sparse View Guided NeRF Update for Object Reconfigurations
by: Lu, Ziqi, et al.
Published: (2024)
by: Lu, Ziqi, et al.
Published: (2024)
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
by: Whitehead, Spencer, et al.
Published: (2024)
by: Whitehead, Spencer, et al.
Published: (2024)
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
by: Karamcheti, Siddharth, et al.
Published: (2024)
by: Karamcheti, Siddharth, et al.
Published: (2024)
First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
by: Ha, Jiwoo, et al.
Published: (2026)
by: Ha, Jiwoo, et al.
Published: (2026)
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
by: Padhi, Trilok, et al.
Published: (2025)
by: Padhi, Trilok, et al.
Published: (2025)
Detecting Korean Food Using Image using Hierarchical Model
by: Lam, Hoang Khanh, et al.
Published: (2024)
by: Lam, Hoang Khanh, et al.
Published: (2024)
Prompt Highlighter: Interactive Control for Multi-Modal LLMs
by: Zhang, Yuechen, et al.
Published: (2023)
by: Zhang, Yuechen, et al.
Published: (2023)
Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
by: Zheng, Haohan, et al.
Published: (2025)
by: Zheng, Haohan, et al.
Published: (2025)
A Quantitative Evaluation of Score Distillation Sampling Based Text-to-3D
by: Fei, Xiaohan, et al.
Published: (2024)
by: Fei, Xiaohan, et al.
Published: (2024)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
by: Geigle, Gregor, et al.
Published: (2024)
by: Geigle, Gregor, et al.
Published: (2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
AI Agents as Universal Task Solvers
by: Achille, Alessandro, et al.
Published: (2025)
by: Achille, Alessandro, et al.
Published: (2025)
Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
by: Pantazopoulos, Georgios, et al.
Published: (2024)
by: Pantazopoulos, Georgios, et al.
Published: (2024)
Similar Items
-
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
by: Trager, Matthew, et al.
Published: (2023) -
Descriminative-Generative Custom Tokens for Vision-Language Models
by: Perera, Pramuditha, et al.
Published: (2025) -
Compositional Structures in Neural Embedding and Interaction Decompositions
by: Trager, Matthew, et al.
Published: (2024) -
PICASO: Permutation-Invariant Context Composition with State Space Models
by: Liu, Tian Yu, et al.
Published: (2025) -
CPR: Retrieval Augmented Generation for Copyright Protection
by: Golatkar, Aditya, et al.
Published: (2024)