MVFNet: Multipurpose Video Forensics Network using Multiple Forms of Forensic Evidence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Tai D., Stamm, Matthew C.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915215021441024
author Nguyen, Tai D.
Stamm, Matthew C.
author_facet Nguyen, Tai D.
Stamm, Matthew C.
contents While videos can be falsified in many different ways, most existing forensic networks are specialized to detect only a single manipulation type (e.g. deepfake, inpainting). This poses a significant issue as the manipulation used to falsify a video is not known a priori. To address this problem, we propose MVFNet - a multipurpose video forensics network capable of detecting multiple types of manipulations including inpainting, deepfakes, splicing, and editing. Our network does this by extracting and jointly analyzing a broad set of forensic feature modalities that capture both spatial and temporal anomalies in falsified videos. To reliably detect and localize fake content of all shapes and sizes, our network employs a novel Multi-Scale Hierarchical Transformer module to identify forensic inconsistencies across multiple spatial scales. Experimental results show that our network obtains state-of-the-art performance in general scenarios where multiple different manipulations are possible, and rivals specialized detectors in targeted scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20991
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MVFNet: Multipurpose Video Forensics Network using Multiple Forms of Forensic Evidence
Nguyen, Tai D.
Stamm, Matthew C.
Computer Vision and Pattern Recognition
While videos can be falsified in many different ways, most existing forensic networks are specialized to detect only a single manipulation type (e.g. deepfake, inpainting). This poses a significant issue as the manipulation used to falsify a video is not known a priori. To address this problem, we propose MVFNet - a multipurpose video forensics network capable of detecting multiple types of manipulations including inpainting, deepfakes, splicing, and editing. Our network does this by extracting and jointly analyzing a broad set of forensic feature modalities that capture both spatial and temporal anomalies in falsified videos. To reliably detect and localize fake content of all shapes and sizes, our network employs a novel Multi-Scale Hierarchical Transformer module to identify forensic inconsistencies across multiple spatial scales. Experimental results show that our network obtains state-of-the-art performance in general scenarios where multiple different manipulations are possible, and rivals specialized detectors in targeted scenarios.
title MVFNet: Multipurpose Video Forensics Network using Multiple Forms of Forensic Evidence
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.20991