Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Xiangrui, Luo, Man, Chatterjee, Agneet, Wei, Hua, Baral, Chitta, Yang, Yezhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
di: Varshney, Neeraj, et al.
Pubblicazione: (2024)
di: Varshney, Neeraj, et al.
Pubblicazione: (2024)
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
di: Fallah, Forouzan, et al.
Pubblicazione: (2025)
di: Fallah, Forouzan, et al.
Pubblicazione: (2025)
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
di: Malaviya, Vatsal, et al.
Pubblicazione: (2025)
di: Malaviya, Vatsal, et al.
Pubblicazione: (2025)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
Dual Caption Preference Optimization for Diffusion Models
di: Saeidi, Amir, et al.
Pubblicazione: (2025)
di: Saeidi, Amir, et al.
Pubblicazione: (2025)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
di: Patel, Maitreya, et al.
Pubblicazione: (2023)
di: Patel, Maitreya, et al.
Pubblicazione: (2023)
Chimera: Compositional Image Generation using Part-based Concepting
di: Singh, Shivam, et al.
Pubblicazione: (2025)
di: Singh, Shivam, et al.
Pubblicazione: (2025)
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
di: Mishra, Venkatesh, et al.
Pubblicazione: (2025)
di: Mishra, Venkatesh, et al.
Pubblicazione: (2025)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
di: Saxon, Michael, et al.
Pubblicazione: (2024)
di: Saxon, Michael, et al.
Pubblicazione: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
di: Yilmaz, Nilay, et al.
Pubblicazione: (2025)
di: Yilmaz, Nilay, et al.
Pubblicazione: (2025)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
di: Chatterjee, Agneet, et al.
Pubblicazione: (2025)
di: Chatterjee, Agneet, et al.
Pubblicazione: (2025)
Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization
di: Saeidi, Amir, et al.
Pubblicazione: (2024)
di: Saeidi, Amir, et al.
Pubblicazione: (2024)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
di: RRV, Aswin, et al.
Pubblicazione: (2024)
di: RRV, Aswin, et al.
Pubblicazione: (2024)
Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective
di: Rajput, Krishna Singh, et al.
Pubblicazione: (2025)
di: Rajput, Krishna Singh, et al.
Pubblicazione: (2025)
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
di: Patel, Nisarg, et al.
Pubblicazione: (2024)
di: Patel, Nisarg, et al.
Pubblicazione: (2024)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
di: Parmar, Mihir, et al.
Pubblicazione: (2024)
di: Parmar, Mihir, et al.
Pubblicazione: (2024)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
di: Luo, Man, et al.
Pubblicazione: (2023)
di: Luo, Man, et al.
Pubblicazione: (2023)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
di: Tyagi, Nemika, et al.
Pubblicazione: (2024)
di: Tyagi, Nemika, et al.
Pubblicazione: (2024)
Map&Make: Schema Guided Text to Table Generation
di: Ahuja, Naman, et al.
Pubblicazione: (2025)
di: Ahuja, Naman, et al.
Pubblicazione: (2025)
The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
di: Varshney, Neeraj, et al.
Pubblicazione: (2023)
di: Varshney, Neeraj, et al.
Pubblicazione: (2023)
PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
di: Parmar, Mihir, et al.
Pubblicazione: (2025)
di: Parmar, Mihir, et al.
Pubblicazione: (2025)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
di: Siingh, Shikhhar, et al.
Pubblicazione: (2025)
di: Siingh, Shikhhar, et al.
Pubblicazione: (2025)
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
di: Saeidi, Amir, et al.
Pubblicazione: (2024)
di: Saeidi, Amir, et al.
Pubblicazione: (2024)
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
di: Pathiraja, Bimsara, et al.
Pubblicazione: (2025)
di: Pathiraja, Bimsara, et al.
Pubblicazione: (2025)
Getting it Right: Improving Spatial Consistency in Text-to-Image Models
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents
di: Uddin, Md Nayem, et al.
Pubblicazione: (2026)
di: Uddin, Md Nayem, et al.
Pubblicazione: (2026)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
di: Parmar, Mihir, et al.
Pubblicazione: (2022)
di: Parmar, Mihir, et al.
Pubblicazione: (2022)
OViP: Online Vision-Language Preference Learning for VLM Hallucination
di: Liu, Shujun, et al.
Pubblicazione: (2025)
di: Liu, Shujun, et al.
Pubblicazione: (2025)
Hallucination Detection and Hallucination Mitigation: An Investigation
di: Luo, Junliang, et al.
Pubblicazione: (2024)
di: Luo, Junliang, et al.
Pubblicazione: (2024)
ThinkTuning: Instilling Cognitive Reflections without Distillation
di: RRV, Aswin, et al.
Pubblicazione: (2025)
di: RRV, Aswin, et al.
Pubblicazione: (2025)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
di: He, Lehan, et al.
Pubblicazione: (2024)
di: He, Lehan, et al.
Pubblicazione: (2024)
Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
di: Kumbhar, Shrinidhi, et al.
Pubblicazione: (2025)
di: Kumbhar, Shrinidhi, et al.
Pubblicazione: (2025)
Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
di: Anantheswaran, Ujjwala, et al.
Pubblicazione: (2024)
di: Anantheswaran, Ujjwala, et al.
Pubblicazione: (2024)
ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
di: Alqurnawi, Yahia, et al.
Pubblicazione: (2026)
di: Alqurnawi, Yahia, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
di: Varshney, Neeraj, et al.
Pubblicazione: (2024) -
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024) -
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
di: Fallah, Forouzan, et al.
Pubblicazione: (2025) -
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024) -
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
di: Malaviya, Vatsal, et al.
Pubblicazione: (2025)