Identifying Predictions That Influence the Future: Detecting Performative Concept Drift in Data Streams

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gower-Winter, Brandon, Krempl, Georg, Dragomiretskiy, Sergey, Jelsma, Tineke, Siebes, Arno
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909560581652480
author Gower-Winter, Brandon
Krempl, Georg
Dragomiretskiy, Sergey
Jelsma, Tineke
Siebes, Arno
author_facet Gower-Winter, Brandon
Krempl, Georg
Dragomiretskiy, Sergey
Jelsma, Tineke
Siebes, Arno
contents Concept Drift has been extensively studied within the context of Stream Learning. However, it is often assumed that the deployed model's predictions play no role in the concept drift the system experiences. Closer inspection reveals that this is not always the case. Automated trading might be prone to self-fulfilling feedback loops. Likewise, malicious entities might adapt to evade detectors in the adversarial setting resulting in a self-negating feedback loop that requires the deployed models to constantly retrain. Such settings where a model may induce concept drift are called performative. In this work, we investigate this phenomenon. Our contributions are as follows: First, we define performative drift within a stream learning setting and distinguish it from other causes of drift. We introduce a novel type of drift detection task, aimed at identifying potential performative concept drift in data streams. We propose a first such performative drift detection approach, called CheckerBoard Performative Drift Detection (CB-PDD). We apply CB-PDD to both synthetic and semi-synthetic datasets that exhibit varying degrees of self-fulfilling feedback loops. Results are positive with CB-PDD showing high efficacy, low false detection rates, resilience to intrinsic drift, comparability to other drift detection techniques, and an ability to effectively detect performative drift in semi-synthetic datasets. Secondly, we highlight the role intrinsic (traditional) drift plays in obfuscating performative drift and discuss the implications of these findings as well as the limitations of CB-PDD.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10545
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Identifying Predictions That Influence the Future: Detecting Performative Concept Drift in Data Streams
Gower-Winter, Brandon
Krempl, Georg
Dragomiretskiy, Sergey
Jelsma, Tineke
Siebes, Arno
Machine Learning
Cryptography and Security
Concept Drift has been extensively studied within the context of Stream Learning. However, it is often assumed that the deployed model's predictions play no role in the concept drift the system experiences. Closer inspection reveals that this is not always the case. Automated trading might be prone to self-fulfilling feedback loops. Likewise, malicious entities might adapt to evade detectors in the adversarial setting resulting in a self-negating feedback loop that requires the deployed models to constantly retrain. Such settings where a model may induce concept drift are called performative. In this work, we investigate this phenomenon. Our contributions are as follows: First, we define performative drift within a stream learning setting and distinguish it from other causes of drift. We introduce a novel type of drift detection task, aimed at identifying potential performative concept drift in data streams. We propose a first such performative drift detection approach, called CheckerBoard Performative Drift Detection (CB-PDD). We apply CB-PDD to both synthetic and semi-synthetic datasets that exhibit varying degrees of self-fulfilling feedback loops. Results are positive with CB-PDD showing high efficacy, low false detection rates, resilience to intrinsic drift, comparability to other drift detection techniques, and an ability to effectively detect performative drift in semi-synthetic datasets. Secondly, we highlight the role intrinsic (traditional) drift plays in obfuscating performative drift and discuss the implications of these findings as well as the limitations of CB-PDD.
title Identifying Predictions That Influence the Future: Detecting Performative Concept Drift in Data Streams
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2412.10545