VidHal: Benchmarking Temporal Hallucinations in Vision LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Choong, Wey Yeh, Guo, Yangyang, Kankanhalli, Mohan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Joint Vision-Language Social Bias Removal for CLIP
di: Zhang, Haoyu, et al.
Pubblicazione: (2024)
di: Zhang, Haoyu, et al.
Pubblicazione: (2024)
SCAN: Bootstrapping Contrastive Pre-training for Data Efficiency
di: Guo, Yangyang, et al.
Pubblicazione: (2024)
di: Guo, Yangyang, et al.
Pubblicazione: (2024)
ELIP: Efficient Discriminative Language-Image Pre-training with Fewer Vision Tokens
di: Guo, Yangyang, et al.
Pubblicazione: (2023)
di: Guo, Yangyang, et al.
Pubblicazione: (2023)
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
di: Saito, Kuniaki, et al.
Pubblicazione: (2026)
di: Saito, Kuniaki, et al.
Pubblicazione: (2026)
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
di: Saito, Kuniaki, et al.
Pubblicazione: (2025)
di: Saito, Kuniaki, et al.
Pubblicazione: (2025)
Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCR
di: Li, Zhenyang, et al.
Pubblicazione: (2024)
di: Li, Zhenyang, et al.
Pubblicazione: (2024)
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
di: Park, Eunkyu, et al.
Pubblicazione: (2025)
di: Park, Eunkyu, et al.
Pubblicazione: (2025)
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
di: Guo, Yangyang, et al.
Pubblicazione: (2024)
di: Guo, Yangyang, et al.
Pubblicazione: (2024)
UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models
di: Guo, Yangyang, et al.
Pubblicazione: (2023)
di: Guo, Yangyang, et al.
Pubblicazione: (2023)
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs
di: Yang, Yiming, et al.
Pubblicazione: (2025)
di: Yang, Yiming, et al.
Pubblicazione: (2025)
Diffusion Facial Forgery Detection
di: Cheng, Harry, et al.
Pubblicazione: (2024)
di: Cheng, Harry, et al.
Pubblicazione: (2024)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
di: Li, Chaoyu, et al.
Pubblicazione: (2024)
di: Li, Chaoyu, et al.
Pubblicazione: (2024)
Word-Anchored Temporal Forgery Localization
di: Wang, Tianyi, et al.
Pubblicazione: (2026)
di: Wang, Tianyi, et al.
Pubblicazione: (2026)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
di: Zhong, Yangyang, et al.
Pubblicazione: (2025)
di: Zhong, Yangyang, et al.
Pubblicazione: (2025)
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
di: Qiu, Han, et al.
Pubblicazione: (2024)
di: Qiu, Han, et al.
Pubblicazione: (2024)
Technical Report for ICML 2024 TiFA Workshop MLLM Attack Challenge: Suffix Injection and Projected Gradient Descent Can Easily Fool An MLLM
di: Guo, Yangyang, et al.
Pubblicazione: (2024)
di: Guo, Yangyang, et al.
Pubblicazione: (2024)
Fair Deepfake Detectors Can Generalize
di: Cheng, Harry, et al.
Pubblicazione: (2025)
di: Cheng, Harry, et al.
Pubblicazione: (2025)
ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding
di: Wang, Xingqi, et al.
Pubblicazione: (2025)
di: Wang, Xingqi, et al.
Pubblicazione: (2025)
EnsemHalDet: Robust VLM Hallucination Detection via Ensemble of Internal State Detectors
di: Miyazato, Ryuhei, et al.
Pubblicazione: (2026)
di: Miyazato, Ryuhei, et al.
Pubblicazione: (2026)
Aggregating Diverse Cue Experts for AI-Generated Image Detection
di: Tan, Lei, et al.
Pubblicazione: (2026)
di: Tan, Lei, et al.
Pubblicazione: (2026)
Enhancing HOI Detection with Contextual Cues from Large Vision-Language Models
di: Zhan, Yu-Wei, et al.
Pubblicazione: (2023)
di: Zhan, Yu-Wei, et al.
Pubblicazione: (2023)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
HalCECE: A Framework for Explainable Hallucination Detection through Conceptual Counterfactuals in Image Captioning
di: Lymperaiou, Maria, et al.
Pubblicazione: (2025)
di: Lymperaiou, Maria, et al.
Pubblicazione: (2025)
FractalForensics: Proactive Deepfake Detection and Localization via Fractal Watermarks
di: Wang, Tianyi, et al.
Pubblicazione: (2025)
di: Wang, Tianyi, et al.
Pubblicazione: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
di: Xu, Yicheng, et al.
Pubblicazione: (2025)
di: Xu, Yicheng, et al.
Pubblicazione: (2025)
Object-Centric Framework for Video Moment Retrieval
di: Li, Zongyao, et al.
Pubblicazione: (2025)
di: Li, Zongyao, et al.
Pubblicazione: (2025)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
di: Cai, Jianfeng, et al.
Pubblicazione: (2025)
di: Cai, Jianfeng, et al.
Pubblicazione: (2025)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
di: Huang, Shiqi, et al.
Pubblicazione: (2026)
di: Huang, Shiqi, et al.
Pubblicazione: (2026)
PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
di: Sun, Fengyuan, et al.
Pubblicazione: (2025)
di: Sun, Fengyuan, et al.
Pubblicazione: (2025)
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
di: Chai, Zenghao, et al.
Pubblicazione: (2024)
di: Chai, Zenghao, et al.
Pubblicazione: (2024)
Learning to Predict Gradients for Semi-Supervised Continual Learning
di: Luo, Yan, et al.
Pubblicazione: (2022)
di: Luo, Yan, et al.
Pubblicazione: (2022)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
di: Chen, Lan, et al.
Pubblicazione: (2025)
di: Chen, Lan, et al.
Pubblicazione: (2025)
MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer
di: Chai, Zenghao, et al.
Pubblicazione: (2025)
di: Chai, Zenghao, et al.
Pubblicazione: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
di: Fang, Ye, et al.
Pubblicazione: (2025)
di: Fang, Ye, et al.
Pubblicazione: (2025)
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
di: Ni, Zhenliang, et al.
Pubblicazione: (2025)
di: Ni, Zhenliang, et al.
Pubblicazione: (2025)
Towards Generalizable Deepfake Detection via Real Distribution Bias Correction
di: Liu, Ming-Hui, et al.
Pubblicazione: (2026)
di: Liu, Ming-Hui, et al.
Pubblicazione: (2026)
E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching
di: Nie, Jiahao, et al.
Pubblicazione: (2026)
di: Nie, Jiahao, et al.
Pubblicazione: (2026)
Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs
di: Nguyen, Dung, et al.
Pubblicazione: (2025)
di: Nguyen, Dung, et al.
Pubblicazione: (2025)
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
di: Yakun, Cui, et al.
Pubblicazione: (2026)
di: Yakun, Cui, et al.
Pubblicazione: (2026)
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
di: Wu, Xiyang, et al.
Pubblicazione: (2024)
di: Wu, Xiyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Joint Vision-Language Social Bias Removal for CLIP
di: Zhang, Haoyu, et al.
Pubblicazione: (2024) -
SCAN: Bootstrapping Contrastive Pre-training for Data Efficiency
di: Guo, Yangyang, et al.
Pubblicazione: (2024) -
ELIP: Efficient Discriminative Language-Image Pre-training with Fewer Vision Tokens
di: Guo, Yangyang, et al.
Pubblicazione: (2023) -
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
di: Saito, Kuniaki, et al.
Pubblicazione: (2026) -
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
di: Saito, Kuniaki, et al.
Pubblicazione: (2025)