Beyond Static Artifacts: A Forensic Benchmark for Video Deepfake Reasoning in Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Zheyuan, Zhao, Qingsong, Wang, Yusong, Huang, Zhaohong, Li, Xinqi, Yuan, Cheng, Shao, Jiaowei, Zhang, Chi, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exposing and Mitigating Temporal Attack in Deepfake Video Detection
by: Gu, Zheyuan, et al.
Published: (2026)
by: Gu, Zheyuan, et al.
Published: (2026)
Enhancing Neural Video Compression of Static Scenes with Positive-Incentive Noise
by: Yuan, Cheng, et al.
Published: (2026)
by: Yuan, Cheng, et al.
Published: (2026)
Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
by: Yuan, Cheng, et al.
Published: (2025)
by: Yuan, Cheng, et al.
Published: (2025)
SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes
by: Leotta, Roberto, et al.
Published: (2026)
by: Leotta, Roberto, et al.
Published: (2026)
HQ-MPSD: A Multilingual Artifact-Controlled Benchmark for Partial Deepfake Speech Detection
by: Li, Menglu, et al.
Published: (2025)
by: Li, Menglu, et al.
Published: (2025)
Forensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language Models
by: Wang, Jin, et al.
Published: (2025)
by: Wang, Jin, et al.
Published: (2025)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Conditional Video Generation for High-Efficiency Video Compression
by: Yi, Fangqiu, et al.
Published: (2025)
by: Yi, Fangqiu, et al.
Published: (2025)
Enhance Vision-Language Alignment with Noise
by: Huang, Sida, et al.
Published: (2024)
by: Huang, Sida, et al.
Published: (2024)
Training-Free Multimodal Deepfake Detection via Graph Reasoning
by: Liu, Yuxin, et al.
Published: (2025)
by: Liu, Yuxin, et al.
Published: (2025)
SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model
by: Cao, Bin, et al.
Published: (2024)
by: Cao, Bin, et al.
Published: (2024)
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
by: Zhang, Da, et al.
Published: (2025)
by: Zhang, Da, et al.
Published: (2025)
Visual Implicit Autoregressive Modeling
by: Jiang, Pengfei, et al.
Published: (2026)
by: Jiang, Pengfei, et al.
Published: (2026)
Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language Models
by: Cai, Wei, et al.
Published: (2025)
by: Cai, Wei, et al.
Published: (2025)
Forensic Similarity for Speech Deepfakes
by: Negroni, Viola, et al.
Published: (2025)
by: Negroni, Viola, et al.
Published: (2025)
BabyVision: Visual Reasoning Beyond Language
by: Chen, Liang, et al.
Published: (2026)
by: Chen, Liang, et al.
Published: (2026)
The Law of Multi-Model Collaboration: Scaling Limits of Model Ensembling for Large Language Models
by: Lu, Dakuan, et al.
Published: (2025)
by: Lu, Dakuan, et al.
Published: (2025)
Single-Pixel Vision-Language Model for Intrinsic Privacy-Preserving Behavioral Intelligence
by: An, Hongjun, et al.
Published: (2026)
by: An, Hongjun, et al.
Published: (2026)
Unleashing Vision-Language Semantics for Deepfake Video Detection
by: Zhu, Jiawen, et al.
Published: (2026)
by: Zhu, Jiawen, et al.
Published: (2026)
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
by: Huang, Xinmiao, et al.
Published: (2025)
by: Huang, Xinmiao, et al.
Published: (2025)
ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning
by: Zhao, Yusong, et al.
Published: (2026)
by: Zhao, Yusong, et al.
Published: (2026)
Media Forensics and Deepfake Systematic Survey
by: CH, Nadeem Jabbar, et al.
Published: (2024)
by: CH, Nadeem Jabbar, et al.
Published: (2024)
VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
by: Xue, Yufei, et al.
Published: (2025)
by: Xue, Yufei, et al.
Published: (2025)
Beyond the Wrapper: Identifying Artifact Reliance in Static Malware Classifiers using TRUSTEE
by: Mohammed, Riyazuddin, et al.
Published: (2026)
by: Mohammed, Riyazuddin, et al.
Published: (2026)
FractalForensics: Proactive Deepfake Detection and Localization via Fractal Watermarks
by: Wang, Tianyi, et al.
Published: (2025)
by: Wang, Tianyi, et al.
Published: (2025)
In Anticipation of Perfect Deepfake: Identity-anchored Artifact-agnostic Detection under Rebalanced Deepfake Detection Protocol
by: Wang, Wei-Han, et al.
Published: (2024)
by: Wang, Wei-Han, et al.
Published: (2024)
INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
by: Bagaria, Anshul
Published: (2025)
by: Bagaria, Anshul
Published: (2025)
AI Flow at the Network Edge
by: Shao, Jiawei, et al.
Published: (2024)
by: Shao, Jiawei, et al.
Published: (2024)
When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
by: Cai, Wei, et al.
Published: (2025)
by: Cai, Wei, et al.
Published: (2025)
FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Pay Less Attention to Deceptive Artifacts: Robust Detection of Compressed Deepfakes on Online Social Networks
by: Li, Manyi, et al.
Published: (2025)
by: Li, Manyi, et al.
Published: (2025)
Prototype-Based Test-Time Adaptation of Vision-Language Models
by: Huang, Zhaohong, et al.
Published: (2026)
by: Huang, Zhaohong, et al.
Published: (2026)
Evidence Packing for Cross-Domain Image Deepfake Detection with LVLMs
by: Liu, Yuxin, et al.
Published: (2026)
by: Liu, Yuxin, et al.
Published: (2026)
Deepfake Forensics Adapter: A Dual-Stream Network for Generalizable Deepfake Detection
by: Liao, Jianfeng, et al.
Published: (2026)
by: Liao, Jianfeng, et al.
Published: (2026)
Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
by: Guo, Longteng, et al.
Published: (2026)
by: Guo, Longteng, et al.
Published: (2026)
Depicting Beyond Scores: Advancing Image Quality Assessment through Multi-modal Language Models
by: You, Zhiyuan, et al.
Published: (2023)
by: You, Zhiyuan, et al.
Published: (2023)
FauForensics: Boosting Audio-Visual Deepfake Detection with Facial Action Units
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Similar Items
-
Exposing and Mitigating Temporal Attack in Deepfake Video Detection
by: Gu, Zheyuan, et al.
Published: (2026) -
Enhancing Neural Video Compression of Static Scenes with Positive-Incentive Noise
by: Yuan, Cheng, et al.
Published: (2026) -
Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
by: Yuan, Cheng, et al.
Published: (2025) -
SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes
by: Leotta, Roberto, et al.
Published: (2026) -
HQ-MPSD: A Multilingual Artifact-Controlled Benchmark for Partial Deepfake Speech Detection
by: Li, Menglu, et al.
Published: (2025)