HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guan, Tianrui, Liu, Fuxiao, Wu, Xiyang, Xian, Ruiqi, Li, Zongxia, Liu, Xiaoyu, Wang, Xijun, Chen, Lichang, Huang, Furong, Yacoob, Yaser, Manocha, Dinesh, Zhou, Tianyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
von: Wu, Xiyang, et al.
Veröffentlicht: (2024)
von: Wu, Xiyang, et al.
Veröffentlicht: (2024)
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
von: Wang, Xijun, et al.
Veröffentlicht: (2023)
von: Wang, Xijun, et al.
Veröffentlicht: (2023)
FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition
von: Xian, Ruiqi, et al.
Veröffentlicht: (2024)
von: Xian, Ruiqi, et al.
Veröffentlicht: (2024)
AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales
von: Guan, Tianrui, et al.
Veröffentlicht: (2024)
von: Guan, Tianrui, et al.
Veröffentlicht: (2024)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
On the Vulnerability of LLM/VLM-Controlled Robotics
von: Wu, Xiyang, et al.
Veröffentlicht: (2024)
von: Wu, Xiyang, et al.
Veröffentlicht: (2024)
DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments
von: Wang, Xijun, et al.
Veröffentlicht: (2024)
von: Wang, Xijun, et al.
Veröffentlicht: (2024)
Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Diffusion Models through Adaptive Attention Modulation
von: Oorloff, Trevine, et al.
Veröffentlicht: (2025)
von: Oorloff, Trevine, et al.
Veröffentlicht: (2025)
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models
von: Wu, Xiyang, et al.
Veröffentlicht: (2026)
von: Wu, Xiyang, et al.
Veröffentlicht: (2026)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
Robot Navigation Using Physically Grounded Vision-Language Models in Outdoor Environments
von: Elnoor, Mohamed, et al.
Veröffentlicht: (2024)
von: Elnoor, Mohamed, et al.
Veröffentlicht: (2024)
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
von: Wu, Xiyang, et al.
Veröffentlicht: (2026)
von: Wu, Xiyang, et al.
Veröffentlicht: (2026)
MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
von: Wu, Xiyang, et al.
Veröffentlicht: (2025)
von: Wu, Xiyang, et al.
Veröffentlicht: (2025)
MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data
von: Li, Zongxia, et al.
Veröffentlicht: (2026)
von: Li, Zongxia, et al.
Veröffentlicht: (2026)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)
GaussianSSC: Triplane-Guided Directional Gaussian Fields for 3D Semantic Completion
von: Xian, Ruiqi, et al.
Veröffentlicht: (2026)
von: Xian, Ruiqi, et al.
Veröffentlicht: (2026)
Large Language Models and Causal Inference in Collaboration: A Survey
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
Paired-CSLiDAR: Height-Stratified Registration for Cross-Source Aerial-Ground LiDAR Pose Refinement
von: Hoover, Montana, et al.
Veröffentlicht: (2026)
von: Hoover, Montana, et al.
Veröffentlicht: (2026)
Towards Understanding In-Context Learning with Contrastive Demonstrations and Saliency Maps
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
First Frame Is the Place to Go for Video Content Customization
von: Chen, Jingxi, et al.
Veröffentlicht: (2025)
von: Chen, Jingxi, et al.
Veröffentlicht: (2025)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
AIDE: Agentically Improve Visual Language Model with Domain Experts
von: Chiu, Ming-Chang, et al.
Veröffentlicht: (2025)
von: Chiu, Ming-Chang, et al.
Veröffentlicht: (2025)
Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models
von: Li, Chaozhuo, et al.
Veröffentlicht: (2025)
von: Li, Chaozhuo, et al.
Veröffentlicht: (2025)
Beyond the Binary
von: Yacoob, Saadia
Veröffentlicht: (2024)
von: Yacoob, Saadia
Veröffentlicht: (2024)
CalibFree: Self-Supervised View Feature Separation for Calibration-Free Multi-Camera Multi-Object Tracking
von: Xian, Ruiqi, et al.
Veröffentlicht: (2026)
von: Xian, Ruiqi, et al.
Veröffentlicht: (2026)
LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
von: Guan, Tianrui, et al.
Veröffentlicht: (2024)
von: Guan, Tianrui, et al.
Veröffentlicht: (2024)
HawkI: Homography & Mutual Information Guidance for 3D-free Single Image to Aerial View
von: Kothandaraman, Divya, et al.
Veröffentlicht: (2023)
von: Kothandaraman, Divya, et al.
Veröffentlicht: (2023)
Forensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language Models
von: Wang, Jin, et al.
Veröffentlicht: (2025)
von: Wang, Jin, et al.
Veröffentlicht: (2025)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
LANCAR: Leveraging Language for Context-Aware Robot Locomotion in Unstructured Environments
von: Shek, Chak Lam, et al.
Veröffentlicht: (2023)
von: Shek, Chak Lam, et al.
Veröffentlicht: (2023)
Self-Rewarding Vision-Language Model via Reasoning Decomposition
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
Exposure of taste buds to potassium permanganate and formalin suppresses the gustatory neural response in the Nile tilapia Oreochromis niloticus (Linnaeus) / Syed Yahiya Yacoob
von: Yahiya Yacoob, Syed
Veröffentlicht: (2002)
von: Yahiya Yacoob, Syed
Veröffentlicht: (2002)
Do Vision-Language Models Understand Compound Nouns?
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
von: Wu, Xiyang, et al.
Veröffentlicht: (2024) -
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
von: Wang, Xijun, et al.
Veröffentlicht: (2023) -
FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition
von: Xian, Ruiqi, et al.
Veröffentlicht: (2024) -
AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales
von: Guan, Tianrui, et al.
Veröffentlicht: (2024) -
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
von: Li, Zongxia, et al.
Veröffentlicht: (2025)