Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lei, Zhanhe, Wang, Zhongyuan, Cheng, Jikang, Huang, Baojin, Yang, Yuhong, Han, Zhen, Liang, Chao, Ye, Dengpan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913149113860096
author Lei, Zhanhe
Wang, Zhongyuan
Cheng, Jikang
Huang, Baojin
Yang, Yuhong
Han, Zhen
Liang, Chao
Ye, Dengpan
author_facet Lei, Zhanhe
Wang, Zhongyuan
Cheng, Jikang
Huang, Baojin
Yang, Yuhong
Han, Zhen
Liang, Chao
Ye, Dengpan
contents Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and generalizable features. In this work, we propose a novel Tutor-Student Reinforcement Learning (TSRL) framework to dynamically optimize the training curriculum. Our method models the training process as a Markov Decision Process where a ``Tutor'' agent learns to guide a ``Student'' (the deepfake detector). The Tutor, implemented as a Proximal Policy Optimization (PPO) agent, observes a rich state representation for each training sample, encapsulating not only its visual features but also its historical learning dynamics, such as EMA loss and forgetting counts. Based on this state, the Tutor takes an action by assigning a continuous weight (0-1) to the sample's loss, thereby dynamically re-weighting the training batch. The Tutor is rewarded based on the Student's immediate performance change, specifically rewarding transitions from incorrect to correct predictions. This strategy encourages the Tutor to learn a curriculum that prioritizes high-value samples, such as hard-but-learnable examples, leading to a more efficient and effective training process. We demonstrate that this adaptive curriculum improves the Student's generalization capabilities against unseen manipulation techniques compared to traditional training methods. Code is available at https://github.com/wannac1/TSRL.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24139
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection
Lei, Zhanhe
Wang, Zhongyuan
Cheng, Jikang
Huang, Baojin
Yang, Yuhong
Han, Zhen
Liang, Chao
Ye, Dengpan
Computer Vision and Pattern Recognition
Machine Learning
Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and generalizable features. In this work, we propose a novel Tutor-Student Reinforcement Learning (TSRL) framework to dynamically optimize the training curriculum. Our method models the training process as a Markov Decision Process where a ``Tutor'' agent learns to guide a ``Student'' (the deepfake detector). The Tutor, implemented as a Proximal Policy Optimization (PPO) agent, observes a rich state representation for each training sample, encapsulating not only its visual features but also its historical learning dynamics, such as EMA loss and forgetting counts. Based on this state, the Tutor takes an action by assigning a continuous weight (0-1) to the sample's loss, thereby dynamically re-weighting the training batch. The Tutor is rewarded based on the Student's immediate performance change, specifically rewarding transitions from incorrect to correct predictions. This strategy encourages the Tutor to learn a curriculum that prioritizes high-value samples, such as hard-but-learnable examples, leading to a more efficient and effective training process. We demonstrate that this adaptive curriculum improves the Student's generalization capabilities against unseen manipulation techniques compared to traditional training methods. Code is available at https://github.com/wannac1/TSRL.
title Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2603.24139