VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Zongxia, Wu, Xiyang, Shi, Guangyao, Qin, Yubin, Du, Hongyang, Liu, Fuxiao, Zhou, Tianyi, Manocha, Dinesh, Boyd-Graber, Jordan Lee |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
par: Li, Zongxia, et autres
Publié: (2025)
par: Li, Zongxia, et autres
Publié: (2025)
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models
par: Wu, Xiyang, et autres
Publié: (2026)
par: Wu, Xiyang, et autres
Publié: (2026)
First Frame Is the Place to Go for Video Content Customization
par: Chen, Jingxi, et autres
Publié: (2025)
par: Chen, Jingxi, et autres
Publié: (2025)
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
par: Wu, Xiyang, et autres
Publié: (2026)
par: Wu, Xiyang, et autres
Publié: (2026)
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
par: Guan, Tianrui, et autres
Publié: (2023)
par: Guan, Tianrui, et autres
Publié: (2023)
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
par: Wu, Xiyang, et autres
Publié: (2024)
par: Wu, Xiyang, et autres
Publié: (2024)
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
par: Wang, Xijun, et autres
Publié: (2023)
par: Wang, Xijun, et autres
Publié: (2023)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
par: Seth, Ashish, et autres
Publié: (2025)
par: Seth, Ashish, et autres
Publié: (2025)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
par: Li, Zongxia, et autres
Publié: (2024)
par: Li, Zongxia, et autres
Publié: (2024)
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
par: Wu, Xiyang, et autres
Publié: (2025)
par: Wu, Xiyang, et autres
Publié: (2025)
Exploring Audio Hallucination in Egocentric Video Understanding
par: Seth, Ashish, et autres
Publié: (2026)
par: Seth, Ashish, et autres
Publié: (2026)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
par: Li, Zongxia, et autres
Publié: (2025)
par: Li, Zongxia, et autres
Publié: (2025)
HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents
par: Jin, Chao, et autres
Publié: (2026)
par: Jin, Chao, et autres
Publié: (2026)
Inst4DGS: Instance-Decomposed 4D Gaussian Splatting with Multi-Video Label Permutation Learning
par: Lee, Yonghan, et autres
Publié: (2026)
par: Lee, Yonghan, et autres
Publié: (2026)
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
par: Li, Zongxia, et autres
Publié: (2024)
par: Li, Zongxia, et autres
Publié: (2024)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
par: Ding, Peng, et autres
Publié: (2024)
par: Ding, Peng, et autres
Publié: (2024)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
par: Seth, Ashish, et autres
Publié: (2024)
par: Seth, Ashish, et autres
Publié: (2024)
HalluCana: Fixing LLM Hallucination with A Canary Lookahead
par: Li, Tianyi, et autres
Publié: (2024)
par: Li, Tianyi, et autres
Publié: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
par: Mondal, Ishani, et autres
Publié: (2024)
par: Mondal, Ishani, et autres
Publié: (2024)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
par: Gu, Feng, et autres
Publié: (2025)
par: Gu, Feng, et autres
Publié: (2025)
Hallucination Mitigation Prompts Long-term Video Understanding
par: Sun, Yiwei, et autres
Publié: (2024)
par: Sun, Yiwei, et autres
Publié: (2024)
HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
par: Yeh, Min-Hsuan, et autres
Publié: (2025)
par: Yeh, Min-Hsuan, et autres
Publié: (2025)
Towards Understanding In-Context Learning with Contrastive Demonstrations and Saliency Maps
par: Liu, Fuxiao, et autres
Publié: (2023)
par: Liu, Fuxiao, et autres
Publié: (2023)
Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition
par: Kothandaraman, Divya, et autres
Publié: (2022)
par: Kothandaraman, Divya, et autres
Publié: (2022)
Self-Rewarding Vision-Language Model via Reasoning Decomposition
par: Li, Zongxia, et autres
Publié: (2025)
par: Li, Zongxia, et autres
Publié: (2025)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
par: Gor, Maharshi, et autres
Publié: (2024)
par: Gor, Maharshi, et autres
Publié: (2024)
MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data
par: Li, Zongxia, et autres
Publié: (2026)
par: Li, Zongxia, et autres
Publié: (2026)
Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of Topic Models
par: Li, Zongxia, et autres
Publié: (2025)
par: Li, Zongxia, et autres
Publié: (2025)
HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
par: Anaokar, Spandan, et autres
Publié: (2025)
par: Anaokar, Spandan, et autres
Publié: (2025)
HalluLens: LLM Hallucination Benchmark
par: Bang, Yejin, et autres
Publié: (2025)
par: Bang, Yejin, et autres
Publié: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
par: Cherif, Ahmed
Publié: (2026)
par: Cherif, Ahmed
Publié: (2026)
SyncTrack4D: Cross-Video Motion Alignment and Video Synchronization for Multi-Video 4D Gaussian Splatting
par: Lee, Yonghan, et autres
Publié: (2025)
par: Lee, Yonghan, et autres
Publié: (2025)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
par: Shu, Matthew, et autres
Publié: (2024)
par: Shu, Matthew, et autres
Publié: (2024)
HalluGen: Synthesizing Realistic and Controllable Hallucinations for Evaluating Image Restoration
par: Kim, Seunghoi, et autres
Publié: (2025)
par: Kim, Seunghoi, et autres
Publié: (2025)
Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills
par: Liu, Dawei, et autres
Publié: (2026)
par: Liu, Dawei, et autres
Publié: (2026)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
par: Srikanth, Neha, et autres
Publié: (2026)
par: Srikanth, Neha, et autres
Publié: (2026)
Labeled Interactive Topic Models
par: Seelman, Kyle, et autres
Publié: (2023)
par: Seelman, Kyle, et autres
Publié: (2023)
"May I Speak?": Multi-modal Attention Guidance in Social VR Group Conversations
par: Lee, Geonsun, et autres
Publié: (2024)
par: Lee, Geonsun, et autres
Publié: (2024)
On the Vulnerability of LLM/VLM-Controlled Robotics
par: Wu, Xiyang, et autres
Publié: (2024)
par: Wu, Xiyang, et autres
Publié: (2024)
Documents similaires
-
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
par: Li, Zongxia, et autres
Publié: (2025) -
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models
par: Wu, Xiyang, et autres
Publié: (2026) -
First Frame Is the Place to Go for Video Content Customization
par: Chen, Jingxi, et autres
Publié: (2025) -
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
par: Wu, Xiyang, et autres
Publié: (2026) -
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
par: Guan, Tianrui, et autres
Publié: (2023)