ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lai, Yingxin, Yu, Zitong, Wang, Jun, Shen, Linlin, Xu, Yong, Cao, Xiaochun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912963534782464
author Lai, Yingxin
Yu, Zitong
Wang, Jun
Shen, Linlin
Xu, Yong
Cao, Xiaochun
author_facet Lai, Yingxin
Yu, Zitong
Wang, Jun
Shen, Linlin
Xu, Yong
Cao, Xiaochun
contents Multimodal Large Language Models (MLLMs) enable interpretable multimedia forensics by generating textual rationales for forgery detection. However, processing dense visual sequences incurs high computational costs, particularly for high-resolution images and videos. Visual token pruning is a practical acceleration strategy, yet existing methods are largely semantic-driven, retaining salient objects while discarding background regions where manipulation traces such as high-frequency anomalies and temporal jitters often reside. To address this issue, we introduce ForensicZip, a training-free framework that reformulates token compression from a forgery-driven perspective. ForensicZip models temporal token evolution as a Birth-Death Optimal Transport problem with a slack dummy node, quantifying physical discontinuities indicating transient generative artifacts. The forensic scoring further integrates transport-based novelty with high-frequency priors to separate forensic evidence from semantic content under large-ratio compression. Experiments on deepfake and AIGC benchmarks show that at 10\% token retention, ForensicZip achieves $2.97\times$ speedup and over 90\% FLOPs reduction while maintaining state-of-the-art detection performance.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12208
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
Lai, Yingxin
Yu, Zitong
Wang, Jun
Shen, Linlin
Xu, Yong
Cao, Xiaochun
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) enable interpretable multimedia forensics by generating textual rationales for forgery detection. However, processing dense visual sequences incurs high computational costs, particularly for high-resolution images and videos. Visual token pruning is a practical acceleration strategy, yet existing methods are largely semantic-driven, retaining salient objects while discarding background regions where manipulation traces such as high-frequency anomalies and temporal jitters often reside. To address this issue, we introduce ForensicZip, a training-free framework that reformulates token compression from a forgery-driven perspective. ForensicZip models temporal token evolution as a Birth-Death Optimal Transport problem with a slack dummy node, quantifying physical discontinuities indicating transient generative artifacts. The forensic scoring further integrates transport-based novelty with high-frequency priors to separate forensic evidence from semantic content under large-ratio compression. Experiments on deepfake and AIGC benchmarks show that at 10\% token retention, ForensicZip achieves $2.97\times$ speedup and over 90\% FLOPs reduction while maintaining state-of-the-art detection performance.
title ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.12208