FAME: A Lightweight Spatio-Temporal Network for Model Attribution of Face-Swap Deepfakes

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ahmad, Wasim, Peng, Yan-Tsung, Chang, Yuan-Hao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913891100917760
author Ahmad, Wasim
Peng, Yan-Tsung
Chang, Yuan-Hao
author_facet Ahmad, Wasim
Peng, Yan-Tsung
Chang, Yuan-Hao
contents The widespread emergence of face-swap Deepfake videos poses growing risks to digital security, privacy, and media integrity, necessitating effective forensic tools for identifying the source of such manipulations. Although most prior research has focused primarily on binary Deepfake detection, the task of model attribution -- determining which generative model produced a given Deepfake -- remains underexplored. In this paper, we introduce FAME (Fake Attribution via Multilevel Embeddings), a lightweight and efficient spatio-temporal framework designed to capture subtle generative artifacts specific to different face-swap models. FAME integrates spatial and temporal attention mechanisms to improve attribution accuracy while remaining computationally efficient. We evaluate our model on three challenging and diverse datasets: Deepfake Detection and Manipulation (DFDM), FaceForensics++, and FakeAVCeleb. Results show that FAME consistently outperforms existing methods in both accuracy and runtime, highlighting its potential for deployment in real-world forensic and information security applications.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FAME: A Lightweight Spatio-Temporal Network for Model Attribution of Face-Swap Deepfakes
Ahmad, Wasim
Peng, Yan-Tsung
Chang, Yuan-Hao
Computer Vision and Pattern Recognition
The widespread emergence of face-swap Deepfake videos poses growing risks to digital security, privacy, and media integrity, necessitating effective forensic tools for identifying the source of such manipulations. Although most prior research has focused primarily on binary Deepfake detection, the task of model attribution -- determining which generative model produced a given Deepfake -- remains underexplored. In this paper, we introduce FAME (Fake Attribution via Multilevel Embeddings), a lightweight and efficient spatio-temporal framework designed to capture subtle generative artifacts specific to different face-swap models. FAME integrates spatial and temporal attention mechanisms to improve attribution accuracy while remaining computationally efficient. We evaluate our model on three challenging and diverse datasets: Deepfake Detection and Manipulation (DFDM), FaceForensics++, and FakeAVCeleb. Results show that FAME consistently outperforms existing methods in both accuracy and runtime, highlighting its potential for deployment in real-world forensic and information security applications.
title FAME: A Lightweight Spatio-Temporal Network for Model Attribution of Face-Swap Deepfakes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.11477