AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Xiyang, Guan, Tianrui, Li, Dianqi, Huang, Shuaiyi, Liu, Xiaoyu, Wang, Xijun, Xian, Ruiqi, Shrivastava, Abhinav, Huang, Furong, Boyd-Graber, Jordan Lee, Zhou, Tianyi, Manocha, Dinesh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
by: Guan, Tianrui, et al.
Published: (2023)
by: Guan, Tianrui, et al.
Published: (2023)
FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition
by: Xian, Ruiqi, et al.
Published: (2024)
by: Xian, Ruiqi, et al.
Published: (2024)
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
by: Wang, Xijun, et al.
Published: (2023)
by: Wang, Xijun, et al.
Published: (2023)
AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales
by: Guan, Tianrui, et al.
Published: (2024)
by: Guan, Tianrui, et al.
Published: (2024)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
by: Wang, Xijun, et al.
Published: (2025)
by: Wang, Xijun, et al.
Published: (2025)
DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments
by: Wang, Xijun, et al.
Published: (2024)
by: Wang, Xijun, et al.
Published: (2024)
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
Paired-CSLiDAR: Height-Stratified Registration for Cross-Source Aerial-Ground LiDAR Pose Refinement
by: Hoover, Montana, et al.
Published: (2026)
by: Hoover, Montana, et al.
Published: (2026)
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2025)
by: Kumar, Pulkit, et al.
Published: (2025)
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations
by: Huang, Shuaiyi, et al.
Published: (2025)
by: Huang, Shuaiyi, et al.
Published: (2025)
Labeled Interactive Topic Models
by: Seelman, Kyle, et al.
Published: (2023)
by: Seelman, Kyle, et al.
Published: (2023)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
by: Gor, Maharshi, et al.
Published: (2024)
by: Gor, Maharshi, et al.
Published: (2024)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
by: Srikanth, Neha, et al.
Published: (2026)
by: Srikanth, Neha, et al.
Published: (2026)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
by: Sung, Yoo Yeon, et al.
Published: (2024)
by: Sung, Yoo Yeon, et al.
Published: (2024)
UVIS: Unsupervised Video Instance Segmentation
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
Mitigating Hallucinations in Diffusion Models through Adaptive Attention Modulation
by: Oorloff, Trevine, et al.
Published: (2025)
by: Oorloff, Trevine, et al.
Published: (2025)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
by: Seth, Ashish, et al.
Published: (2024)
by: Seth, Ashish, et al.
Published: (2024)
ARDuP: Active Region Video Diffusion for Universal Policies
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
What is Point Supervision Worth in Video Instance Segmentation?
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models
by: Lee, Seungjae, et al.
Published: (2025)
by: Lee, Seungjae, et al.
Published: (2025)
Efficient Continuous Video Flow Model for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
by: Shu, Matthew, et al.
Published: (2024)
by: Shu, Matthew, et al.
Published: (2024)
CalibFree: Self-Supervised View Feature Separation for Calibration-Free Multi-Camera Multi-Object Tracking
by: Xian, Ruiqi, et al.
Published: (2026)
by: Xian, Ruiqi, et al.
Published: (2026)
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
by: Chen, Dongping, et al.
Published: (2026)
by: Chen, Dongping, et al.
Published: (2026)
HawkI: Homography & Mutual Information Guidance for 3D-free Single Image to Aerial View
by: Kothandaraman, Divya, et al.
Published: (2023)
by: Kothandaraman, Divya, et al.
Published: (2023)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
by: Wang, Xijun, et al.
Published: (2025)
by: Wang, Xijun, et al.
Published: (2025)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
by: Gor, Maharshi, et al.
Published: (2026)
by: Gor, Maharshi, et al.
Published: (2026)
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
by: Hoyle, Alexander, et al.
Published: (2025)
by: Hoyle, Alexander, et al.
Published: (2025)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Inst4DGS: Instance-Decomposed 4D Gaussian Splatting with Multi-Video Label Permutation Learning
by: Lee, Yonghan, et al.
Published: (2026)
by: Lee, Yonghan, et al.
Published: (2026)
Transfer Q Star: Principled Decoding for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
V-Trans4Style: Visual Transition Recommendation for Video Production Style Adaptation
by: Guhan, Pooja, et al.
Published: (2025)
by: Guhan, Pooja, et al.
Published: (2025)
PACE: Data-Driven Virtual Agent Interaction in Dense and Cluttered Environments
by: Mullen, James, et al.
Published: (2023)
by: Mullen, James, et al.
Published: (2023)
LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
by: Guan, Tianrui, et al.
Published: (2024)
by: Guan, Tianrui, et al.
Published: (2024)
Similar Items
-
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
by: Guan, Tianrui, et al.
Published: (2023) -
FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition
by: Xian, Ruiqi, et al.
Published: (2024) -
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
by: Wang, Xijun, et al.
Published: (2023) -
AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales
by: Guan, Tianrui, et al.
Published: (2024) -
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
by: Li, Zongxia, et al.
Published: (2025)