Towards multi-modal forgery representation learning for AI-generated video detection and localization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Le, Dat, Nguyen, Khoa, Wang, Xin, Hu, Shu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909025728200704
author Le, Dat
Nguyen, Khoa
Wang, Xin
Hu, Shu
author_facet Le, Dat
Nguyen, Khoa
Wang, Xin
Hu, Shu
contents Recent advances in generative AI have democratized video creation at scale. AI-generated videos, including partially manipulated clips across visual and audio channels, pose escalating risks of semantic distortion and misuse, which motivates the need for reliable detection tools. Most existing AI-generated video detectors remain limited by single- or partial-modality of data modeling and the lack of fine-grained temporal forgery localization. To address these challenges, our primary novelty introduces a core architecture that jointly integrates an LMM semantic branch with a spatio-temporal (ST) visual branch and a multi-scale partial-spoof (PS) audio branch. This multi-modal approach enables simultaneous detection and fine-grained temporal localization of partially manipulated AI-generated video forgeries. Extensive experiments show that this approach outperforms existing state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07232
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards multi-modal forgery representation learning for AI-generated video detection and localization
Le, Dat
Nguyen, Khoa
Wang, Xin
Hu, Shu
Computer Vision and Pattern Recognition
Recent advances in generative AI have democratized video creation at scale. AI-generated videos, including partially manipulated clips across visual and audio channels, pose escalating risks of semantic distortion and misuse, which motivates the need for reliable detection tools. Most existing AI-generated video detectors remain limited by single- or partial-modality of data modeling and the lack of fine-grained temporal forgery localization. To address these challenges, our primary novelty introduces a core architecture that jointly integrates an LMM semantic branch with a spatio-temporal (ST) visual branch and a multi-scale partial-spoof (PS) audio branch. This multi-modal approach enables simultaneous detection and fine-grained temporal localization of partially manipulated AI-generated video forgeries. Extensive experiments show that this approach outperforms existing state-of-the-art methods.
title Towards multi-modal forgery representation learning for AI-generated video detection and localization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.07232