Fair-FLIP: Fair Deepfake Detection with Fairness-Oriented Final Layer Input Prioritising

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Szandala, Tomasz, Ezzeddine, Fatima, Rusin, Natalia, Giordano, Silvia, Ayoub, Omran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908447402885120
author Szandala, Tomasz
Ezzeddine, Fatima
Rusin, Natalia
Giordano, Silvia
Ayoub, Omran
author_facet Szandala, Tomasz
Ezzeddine, Fatima
Rusin, Natalia
Giordano, Silvia
Ayoub, Omran
contents Artificial Intelligence-generated content has become increasingly popular, yet its malicious use, particularly the deepfakes, poses a serious threat to public trust and discourse. While deepfake detection methods achieve high predictive performance, they often exhibit biases across demographic attributes such as ethnicity and gender. In this work, we tackle the challenge of fair deepfake detection, aiming to mitigate these biases while maintaining robust detection capabilities. To this end, we propose a novel post-processing approach, referred to as Fairness-Oriented Final Layer Input Prioritising (Fair-FLIP), that reweights a trained model's final-layer inputs to reduce subgroup disparities, prioritising those with low variability while demoting highly variable ones. Experimental results comparing Fair-FLIP to both the baseline (without fairness-oriented de-biasing) and state-of-the-art approaches show that Fair-FLIP can enhance fairness metrics by up to 30% while maintaining baseline accuracy, with only a negligible reduction of 0.25%. Code is available on Github: https://github.com/szandala/fair-deepfake-detection-toolbox
format Preprint
id arxiv_https___arxiv_org_abs_2507_08912
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fair-FLIP: Fair Deepfake Detection with Fairness-Oriented Final Layer Input Prioritising
Szandala, Tomasz
Ezzeddine, Fatima
Rusin, Natalia
Giordano, Silvia
Ayoub, Omran
Machine Learning
Artificial Intelligence
Artificial Intelligence-generated content has become increasingly popular, yet its malicious use, particularly the deepfakes, poses a serious threat to public trust and discourse. While deepfake detection methods achieve high predictive performance, they often exhibit biases across demographic attributes such as ethnicity and gender. In this work, we tackle the challenge of fair deepfake detection, aiming to mitigate these biases while maintaining robust detection capabilities. To this end, we propose a novel post-processing approach, referred to as Fairness-Oriented Final Layer Input Prioritising (Fair-FLIP), that reweights a trained model's final-layer inputs to reduce subgroup disparities, prioritising those with low variability while demoting highly variable ones. Experimental results comparing Fair-FLIP to both the baseline (without fairness-oriented de-biasing) and state-of-the-art approaches show that Fair-FLIP can enhance fairness metrics by up to 30% while maintaining baseline accuracy, with only a negligible reduction of 0.25%. Code is available on Github: https://github.com/szandala/fair-deepfake-detection-toolbox
title Fair-FLIP: Fair Deepfake Detection with Fairness-Oriented Final Layer Input Prioritising
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2507.08912