DMRL: Data- and Model-aware Reward Learning for Data Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhiqiang, Cheng, Ruoxi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913829254856704
author Wang, Zhiqiang
Cheng, Ruoxi
author_facet Wang, Zhiqiang
Cheng, Ruoxi
contents Large language models (LLMs) are inherently vulnerable to unintended privacy breaches. Consequently, systematic red-teaming research is essential for developing robust defense mechanisms. However, current data extraction methods suffer from several limitations: (1) rely on dataset duplicates (addressable via deduplication), (2) depend on prompt engineering (now countered by detection and defense), and (3) rely on random-search adversarial generation. To address these challenges, we propose DMRL, a Data- and Model-aware Reward Learning approach for data extraction. This technique leverages inverse reinforcement learning to extract sensitive data from LLMs. Our method consists of two main components: (1) constructing an introspective reasoning dataset that captures leakage mindsets to guide model behavior, and (2) training reward models with Group Relative Policy Optimization (GRPO), dynamically tuning optimization based on task difficulty at both the data and model levels. Comprehensive experiments across various LLMs demonstrate that DMRL outperforms all baseline methods in data extraction performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06284
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DMRL: Data- and Model-aware Reward Learning for Data Extraction
Wang, Zhiqiang
Cheng, Ruoxi
Machine Learning
Cryptography and Security
Large language models (LLMs) are inherently vulnerable to unintended privacy breaches. Consequently, systematic red-teaming research is essential for developing robust defense mechanisms. However, current data extraction methods suffer from several limitations: (1) rely on dataset duplicates (addressable via deduplication), (2) depend on prompt engineering (now countered by detection and defense), and (3) rely on random-search adversarial generation. To address these challenges, we propose DMRL, a Data- and Model-aware Reward Learning approach for data extraction. This technique leverages inverse reinforcement learning to extract sensitive data from LLMs. Our method consists of two main components: (1) constructing an introspective reasoning dataset that captures leakage mindsets to guide model behavior, and (2) training reward models with Group Relative Policy Optimization (GRPO), dynamically tuning optimization based on task difficulty at both the data and model levels. Comprehensive experiments across various LLMs demonstrate that DMRL outperforms all baseline methods in data extraction performance.
title DMRL: Data- and Model-aware Reward Learning for Data Extraction
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2505.06284