Keeping up with dynamic attackers: Certifying robustness to adaptive online data poisoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bose, Avinandan, Lessard, Laurent, Fazel, Maryam, Dvijotham, Krishnamurthy Dj
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913704660959232
author Bose, Avinandan
Lessard, Laurent
Fazel, Maryam
Dvijotham, Krishnamurthy Dj
author_facet Bose, Avinandan
Lessard, Laurent
Fazel, Maryam
Dvijotham, Krishnamurthy Dj
contents The rise of foundation models fine-tuned on human feedback from potentially untrusted users has increased the risk of adversarial data poisoning, necessitating the study of robustness of learning algorithms against such attacks. Existing research on provable certified robustness against data poisoning attacks primarily focuses on certifying robustness for static adversaries who modify a fraction of the dataset used to train the model before the training algorithm is applied. In practice, particularly when learning from human feedback in an online sense, adversaries can observe and react to the learning process and inject poisoned samples that optimize adversarial objectives better than when they are restricted to poisoning a static dataset once, before the learning algorithm is applied. Indeed, it has been shown in prior work that online dynamic adversaries can be significantly more powerful than static ones. We present a novel framework for computing certified bounds on the impact of dynamic poisoning, and use these certificates to design robust learning algorithms. We give an illustration of the framework for the mean estimation and binary classification problems and outline directions for extending this in further work. The code to implement our certificates and replicate our results is available at https://github.com/Avinandan22/Certified-Robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16737
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Keeping up with dynamic attackers: Certifying robustness to adaptive online data poisoning
Bose, Avinandan
Lessard, Laurent
Fazel, Maryam
Dvijotham, Krishnamurthy Dj
Machine Learning
The rise of foundation models fine-tuned on human feedback from potentially untrusted users has increased the risk of adversarial data poisoning, necessitating the study of robustness of learning algorithms against such attacks. Existing research on provable certified robustness against data poisoning attacks primarily focuses on certifying robustness for static adversaries who modify a fraction of the dataset used to train the model before the training algorithm is applied. In practice, particularly when learning from human feedback in an online sense, adversaries can observe and react to the learning process and inject poisoned samples that optimize adversarial objectives better than when they are restricted to poisoning a static dataset once, before the learning algorithm is applied. Indeed, it has been shown in prior work that online dynamic adversaries can be significantly more powerful than static ones. We present a novel framework for computing certified bounds on the impact of dynamic poisoning, and use these certificates to design robust learning algorithms. We give an illustration of the framework for the mean estimation and binary classification problems and outline directions for extending this in further work. The code to implement our certificates and replicate our results is available at https://github.com/Avinandan22/Certified-Robustness.
title Keeping up with dynamic attackers: Certifying robustness to adaptive online data poisoning
topic Machine Learning
url https://arxiv.org/abs/2502.16737