Hallucination Mitigation Prompts Long-term Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Yiwei, Liu, Zhihang, Liu, Chuanbin, Pu, Bowei, Zhang, Zhihan, Xie, Hongtao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
by: Pu, Bowei, et al.
Published: (2025)
by: Pu, Bowei, et al.
Published: (2025)
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
by: Li, Yinglu, et al.
Published: (2025)
by: Li, Yinglu, et al.
Published: (2025)
From Evaluation to Defense: Advancing Safety in Video Large Language Models
by: Sun, Yiwei, et al.
Published: (2025)
by: Sun, Yiwei, et al.
Published: (2025)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
by: Wang, Jiankang, et al.
Published: (2025)
by: Wang, Jiankang, et al.
Published: (2025)
Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
by: Liu, Zhihang, et al.
Published: (2025)
by: Liu, Zhihang, et al.
Published: (2025)
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
by: Sun, Yiming, et al.
Published: (2025)
by: Sun, Yiming, et al.
Published: (2025)
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
by: Gao, Hongcheng, et al.
Published: (2025)
by: Gao, Hongcheng, et al.
Published: (2025)
PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning
by: Ding, Xinpeng, et al.
Published: (2025)
by: Ding, Xinpeng, et al.
Published: (2025)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models
by: Liu, Zijian, et al.
Published: (2026)
by: Liu, Zijian, et al.
Published: (2026)
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
by: Ma, Ji, et al.
Published: (2026)
by: Ma, Ji, et al.
Published: (2026)
PROBE: Diagnosing Residual Concept Capacity in Erased Text-to-Video Diffusion Models
by: Xie, Yiwei, et al.
Published: (2026)
by: Xie, Yiwei, et al.
Published: (2026)
PosterMaker: Towards High-Quality Product Poster Generation with Accurate Text Rendering
by: Gao, Yifan, et al.
Published: (2025)
by: Gao, Yifan, et al.
Published: (2025)
Where Concept Erasure Should Occur: Concept-Layer Alignment in Text-to-Video Diffusion Models
by: Xie, Yiwei, et al.
Published: (2026)
by: Xie, Yiwei, et al.
Published: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
by: Fu, Honghao, et al.
Published: (2026)
by: Fu, Honghao, et al.
Published: (2026)
CoS: Chain-of-Shot Prompting for Long Video Understanding
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
by: Qiu, Chenhao, et al.
Published: (2026)
by: Qiu, Chenhao, et al.
Published: (2026)
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
by: Zhang, Huaxin, et al.
Published: (2024)
by: Zhang, Huaxin, et al.
Published: (2024)
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
by: Lu, Hao, et al.
Published: (2025)
by: Lu, Hao, et al.
Published: (2025)
Multi-Sentence Grounding for Long-term Instructional Video
by: Li, Zeqian, et al.
Published: (2023)
by: Li, Zeqian, et al.
Published: (2023)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
by: Xue, Zhucun, et al.
Published: (2025)
by: Xue, Zhucun, et al.
Published: (2025)
Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting
by: Wu, Jiarui, et al.
Published: (2025)
by: Wu, Jiarui, et al.
Published: (2025)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
by: Sun, Shangkun, et al.
Published: (2024)
by: Sun, Shangkun, et al.
Published: (2024)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
by: Cai, Jianfeng, et al.
Published: (2025)
by: Cai, Jianfeng, et al.
Published: (2025)
Energy-Guided Decoding for Object Hallucination Mitigation
by: Liu, Xixi, et al.
Published: (2025)
by: Liu, Xixi, et al.
Published: (2025)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
by: Wu, Yike, et al.
Published: (2025)
by: Wu, Yike, et al.
Published: (2025)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
by: Qiu, Jihao, et al.
Published: (2026)
by: Qiu, Jihao, et al.
Published: (2026)
Long Video Understanding with Learnable Retrieval in Video-Language Models
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
by: Xie, Yuan, et al.
Published: (2025)
by: Xie, Yuan, et al.
Published: (2025)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
Multi-granularity Correspondence Learning from Long-term Noisy Videos
by: Lin, Yijie, et al.
Published: (2024)
by: Lin, Yijie, et al.
Published: (2024)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
by: Lin, Jingyang, et al.
Published: (2025)
by: Lin, Jingyang, et al.
Published: (2025)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
by: Jin, Hongbo, et al.
Published: (2025)
by: Jin, Hongbo, et al.
Published: (2025)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
by: Zheng, Haojie, et al.
Published: (2024)
by: Zheng, Haojie, et al.
Published: (2024)
Video World Models with Long-term Spatial Memory
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Similar Items
-
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
by: Pu, Bowei, et al.
Published: (2025) -
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
by: Li, Yinglu, et al.
Published: (2025) -
From Evaluation to Defense: Advancing Safety in Video Large Language Models
by: Sun, Yiwei, et al.
Published: (2025) -
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
by: Wang, Jiankang, et al.
Published: (2025) -
Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
by: Liu, Zhihang, et al.
Published: (2025)