Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sasnauskas, Paulius, Yalın, Yiğit, Radanović, Goran
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909809959239680
author Sasnauskas, Paulius
Yalın, Yiğit
Radanović, Goran
author_facet Sasnauskas, Paulius
Yalın, Yiğit
Radanović, Goran
contents We study the corruption-robustness of in-context reinforcement learning (ICRL), focusing on the Decision-Pretrained Transformer (DPT, Lee et al., 2023). To address the challenge of reward poisoning attacks targeting the DPT, we propose a novel adversarial training framework, called Adversarially Trained Decision-Pretrained Transformer (AT-DPT). Our method simultaneously trains an attacker to minimize the true reward of the DPT by poisoning environment rewards, and a DPT model to infer optimal actions from the poisoned data. We evaluate the effectiveness of our approach against standard bandit algorithms, including robust baselines designed to handle reward contamination. Our results show that the proposed method significantly outperforms these baselines in bandit settings, under a learned attacker. We additionally evaluate AT-DPT on an adaptive attacker, and observe similar results. Furthermore, we extend our evaluation to the MDP setting, confirming that the robustness observed in bandit scenarios generalizes to more complex environments.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06891
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
Sasnauskas, Paulius
Yalın, Yiğit
Radanović, Goran
Machine Learning
Cryptography and Security
We study the corruption-robustness of in-context reinforcement learning (ICRL), focusing on the Decision-Pretrained Transformer (DPT, Lee et al., 2023). To address the challenge of reward poisoning attacks targeting the DPT, we propose a novel adversarial training framework, called Adversarially Trained Decision-Pretrained Transformer (AT-DPT). Our method simultaneously trains an attacker to minimize the true reward of the DPT by poisoning environment rewards, and a DPT model to infer optimal actions from the poisoned data. We evaluate the effectiveness of our approach against standard bandit algorithms, including robust baselines designed to handle reward contamination. Our results show that the proposed method significantly outperforms these baselines in bandit settings, under a learned attacker. We additionally evaluate AT-DPT on an adaptive attacker, and observe similar results. Furthermore, we extend our evaluation to the MDP setting, confirming that the robustness observed in bandit scenarios generalizes to more complex environments.
title Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2506.06891