Saved in:
| Main Authors: | Wu, Tsung-Han, Lee, Heekyung, Ge, Jiaxin, Gonzalez, Joseph E., Darrell, Trevor, Chan, David M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.13169 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025)
by: Lee, Heekyung, et al.
Published: (2025)
Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
by: Wu, Tsung-Han, et al.
Published: (2024)
by: Wu, Tsung-Han, et al.
Published: (2024)
Constantly Improving Image Models Need Constantly Improving Benchmarks
by: Ge, Jiaxin, et al.
Published: (2025)
by: Ge, Jiaxin, et al.
Published: (2025)
ALOHa: A New Measure for Hallucination in Captioning Models
by: Petryk, Suzanne, et al.
Published: (2024)
by: Petryk, Suzanne, et al.
Published: (2024)
REOrdering Patches Improves Vision Models
by: Kutscher, Declan, et al.
Published: (2025)
by: Kutscher, Declan, et al.
Published: (2025)
Vision-Language Models Create Cross-Modal Task Representations
by: Luo, Grace, et al.
Published: (2024)
by: Luo, Grace, et al.
Published: (2024)
CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
by: Wu, Tsung-Han, et al.
Published: (2024)
by: Wu, Tsung-Han, et al.
Published: (2024)
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023)
by: Ge, Jiaxin, et al.
Published: (2023)
When Do We Not Need Larger Vision Models?
by: Shi, Baifeng, et al.
Published: (2024)
by: Shi, Baifeng, et al.
Published: (2024)
Hallucination Elimination and Semantic Enhancement Framework for Vision-Language Models in Traffic Scenarios
by: Fan, Jiaqi, et al.
Published: (2024)
by: Fan, Jiaqi, et al.
Published: (2024)
Discovering Divergent Representations between Text-to-Image Models
by: Dunlap, Lisa, et al.
Published: (2025)
by: Dunlap, Lisa, et al.
Published: (2025)
Dual-Pathway Circuits of Object Hallucination in Vision-Language Models
by: Liu, Jiaxin, et al.
Published: (2026)
by: Liu, Jiaxin, et al.
Published: (2026)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
by: Lian, Long, et al.
Published: (2023)
by: Lian, Long, et al.
Published: (2023)
Resampling Benchmark for Efficient Comprehensive Evaluation of Large Vision-Language Models
by: Suzuki, Teppei, et al.
Published: (2025)
by: Suzuki, Teppei, et al.
Published: (2025)
VisionArena: 230K Real World User-VLM Conversations with Preference Labels
by: Chou, Christopher, et al.
Published: (2024)
by: Chou, Christopher, et al.
Published: (2024)
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
by: Yin, Shaofeng, et al.
Published: (2026)
by: Yin, Shaofeng, et al.
Published: (2026)
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
by: Wu, Xiyang, et al.
Published: (2024)
by: Wu, Xiyang, et al.
Published: (2024)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
by: Geigle, Gregor, et al.
Published: (2024)
by: Geigle, Gregor, et al.
Published: (2024)
Self-Consistency as a Free Lunch: Reducing Hallucinations in Vision-Language Models via Self-Reflection
by: Han, Mingfei, et al.
Published: (2025)
by: Han, Mingfei, et al.
Published: (2025)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
by: Song, Jiale, et al.
Published: (2026)
by: Song, Jiale, et al.
Published: (2026)
Skip \n: A Simple Method to Reduce Hallucination in Large Vision-Language Models
by: Han, Zongbo, et al.
Published: (2024)
by: Han, Zongbo, et al.
Published: (2024)
MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
by: Zhao, Qiyan, et al.
Published: (2025)
by: Zhao, Qiyan, et al.
Published: (2025)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
by: Liu, Sheng, et al.
Published: (2024)
by: Liu, Sheng, et al.
Published: (2024)
Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
by: Wu, Mingrui, et al.
Published: (2024)
by: Wu, Mingrui, et al.
Published: (2024)
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
by: Agrawal, Aakriti, et al.
Published: (2025)
by: Agrawal, Aakriti, et al.
Published: (2025)
Analyzing The Language of Visual Tokens
by: Chan, David M., et al.
Published: (2024)
by: Chan, David M., et al.
Published: (2024)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token
by: Kogilathota, Sai Akhil, et al.
Published: (2026)
by: Kogilathota, Sai Akhil, et al.
Published: (2026)
Pose Priors from Language Models
by: Subramanian, Sanjay, et al.
Published: (2024)
by: Subramanian, Sanjay, et al.
Published: (2024)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
by: Huang, Brandon, et al.
Published: (2025)
by: Huang, Brandon, et al.
Published: (2025)
Collaborative Multi-Mode Pruning for Vision-Language Models
by: Wu, Zimeng, et al.
Published: (2026)
by: Wu, Zimeng, et al.
Published: (2026)
Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
The Role of Background Information in Reducing Object Hallucination in Vision-Language Models: Insights from Cutoff API Prompting
by: Tomita, Masayo, et al.
Published: (2025)
by: Tomita, Masayo, et al.
Published: (2025)
Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
by: Lyu, Xinyu, et al.
Published: (2024)
by: Lyu, Xinyu, et al.
Published: (2024)
From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models
by: Gong, Weile, et al.
Published: (2026)
by: Gong, Weile, et al.
Published: (2026)
Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models
by: Sanogo, Kassoum, et al.
Published: (2025)
by: Sanogo, Kassoum, et al.
Published: (2025)
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
by: Han, Yuhang, et al.
Published: (2026)
by: Han, Yuhang, et al.
Published: (2026)
Revealing Multi-View Hallucination in Large Vision-Language Models
by: Park, Wooje, et al.
Published: (2026)
by: Park, Wooje, et al.
Published: (2026)
Vector Quantized Feature Fields for Fast 3D Semantic Lifting
by: Tang, George, et al.
Published: (2025)
by: Tang, George, et al.
Published: (2025)
Dual-Process Image Generation
by: Luo, Grace, et al.
Published: (2025)
by: Luo, Grace, et al.
Published: (2025)
Similar Items
-
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025) -
Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
by: Wu, Tsung-Han, et al.
Published: (2024) -
Constantly Improving Image Models Need Constantly Improving Benchmarks
by: Ge, Jiaxin, et al.
Published: (2025) -
ALOHa: A New Measure for Hallucination in Captioning Models
by: Petryk, Suzanne, et al.
Published: (2024) -
REOrdering Patches Improves Vision Models
by: Kutscher, Declan, et al.
Published: (2025)