Potent but Stealthy: Rethink Profile Pollution against Sequential Recommendation via Bi-level Constrained Reinforcement Paradigm

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Jiajie, Nan, Zihan, Ma, Yunshan, Xia, Xiaobo, Feng, Xiaohua, Liu, Weiming, Chen, Xiang, Zheng, Xiaolin, Chen, Chaochao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915689548218368
author Su, Jiajie
Nan, Zihan
Ma, Yunshan
Xia, Xiaobo
Feng, Xiaohua
Liu, Weiming
Chen, Xiang
Zheng, Xiaolin
Chen, Chaochao
author_facet Su, Jiajie
Nan, Zihan
Ma, Yunshan
Xia, Xiaobo
Feng, Xiaohua
Liu, Weiming
Chen, Xiang
Zheng, Xiaolin
Chen, Chaochao
contents Sequential Recommenders, which exploit dynamic user intents through interaction sequences, is vulnerable to adversarial attacks. While existing attacks primarily rely on data poisoning, they require large-scale user access or fake profiles thus lacking practicality. In this paper, we focus on the Profile Pollution Attack that subtly contaminates partial user interactions to induce targeted mispredictions. Previous PPA methods suffer from two limitations, i.e., i) over-reliance on sequence horizon impact restricts fine-grained perturbations on item transitions, and ii) holistic modifications cause detectable distribution shifts. To address these challenges, we propose a constrained reinforcement driven attack CREAT that synergizes a bi-level optimization framework with multi-reward reinforcement learning to balance adversarial efficacy and stealthiness. We first develop a Pattern Balanced Rewarding Policy, which integrates pattern inversion rewards to invert critical patterns and distribution consistency rewards to minimize detectable shifts via unbalanced co-optimal transport. Then we employ a Constrained Group Relative Reinforcement Learning paradigm, enabling step-wise perturbations through dynamic barrier constraints and group-shared experience replay, achieving targeted pollution with minimal detectability. Extensive experiments demonstrate the effectiveness of CREAT.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09392
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Potent but Stealthy: Rethink Profile Pollution against Sequential Recommendation via Bi-level Constrained Reinforcement Paradigm
Su, Jiajie
Nan, Zihan
Ma, Yunshan
Xia, Xiaobo
Feng, Xiaohua
Liu, Weiming
Chen, Xiang
Zheng, Xiaolin
Chen, Chaochao
Machine Learning
Artificial Intelligence
Sequential Recommenders, which exploit dynamic user intents through interaction sequences, is vulnerable to adversarial attacks. While existing attacks primarily rely on data poisoning, they require large-scale user access or fake profiles thus lacking practicality. In this paper, we focus on the Profile Pollution Attack that subtly contaminates partial user interactions to induce targeted mispredictions. Previous PPA methods suffer from two limitations, i.e., i) over-reliance on sequence horizon impact restricts fine-grained perturbations on item transitions, and ii) holistic modifications cause detectable distribution shifts. To address these challenges, we propose a constrained reinforcement driven attack CREAT that synergizes a bi-level optimization framework with multi-reward reinforcement learning to balance adversarial efficacy and stealthiness. We first develop a Pattern Balanced Rewarding Policy, which integrates pattern inversion rewards to invert critical patterns and distribution consistency rewards to minimize detectable shifts via unbalanced co-optimal transport. Then we employ a Constrained Group Relative Reinforcement Learning paradigm, enabling step-wise perturbations through dynamic barrier constraints and group-shared experience replay, achieving targeted pollution with minimal detectability. Extensive experiments demonstrate the effectiveness of CREAT.
title Potent but Stealthy: Rethink Profile Pollution against Sequential Recommendation via Bi-level Constrained Reinforcement Paradigm
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.09392