ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liang, Jiacheng, Ma, Yao, Kumarage, Tharindu, Krishna, Satyapriya, Gupta, Rahul, Chang, Kai-Wei, Galstyan, Aram, Peris, Charith
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917423695790080
author Liang, Jiacheng
Ma, Yao
Kumarage, Tharindu
Krishna, Satyapriya
Gupta, Rahul
Chang, Kai-Wei
Galstyan, Aram
Peris, Charith
author_facet Liang, Jiacheng
Ma, Yao
Kumarage, Tharindu
Krishna, Satyapriya
Gupta, Rahul
Chang, Kai-Wei
Galstyan, Aram
Peris, Charith
contents Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an imperfect Reward Model (RM) can become a single point of failure when it fails to penalize unsafe behaviors. While existing red-teaming approaches primarily target policy-level weaknesses, they overlook what we term systemic weaknesses cases where both the core LLM and the RM fail in tandem. We present ARES, a framework that systematically discovers and mitigates such dual vulnerabilities. ARES employs a ``Safety Mentor'' that dynamically composes semantically coherent adversarial prompts by combining structured component types (topics, personas, tactics, goals) and generates corresponding malicious and safe responses. This dual-targeting approach exposes weaknesses in both the core LLM and the RM simultaneously. Using the vulnerabilities gained, ARES implements a two-stage repair process: first fine-tuning the RM to better detect harmful content, then leveraging the improved RM to optimize the core model. Experiments across multiple adversarial safety benchmarks demonstrate that ARES substantially enhances safety robustness while preserving model capabilities, establishing a new paradigm for comprehensive RLHF safety alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18789
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
Liang, Jiacheng
Ma, Yao
Kumarage, Tharindu
Krishna, Satyapriya
Gupta, Rahul
Chang, Kai-Wei
Galstyan, Aram
Peris, Charith
Artificial Intelligence
Cryptography and Security
Machine Learning
Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an imperfect Reward Model (RM) can become a single point of failure when it fails to penalize unsafe behaviors. While existing red-teaming approaches primarily target policy-level weaknesses, they overlook what we term systemic weaknesses cases where both the core LLM and the RM fail in tandem. We present ARES, a framework that systematically discovers and mitigates such dual vulnerabilities. ARES employs a ``Safety Mentor'' that dynamically composes semantically coherent adversarial prompts by combining structured component types (topics, personas, tactics, goals) and generates corresponding malicious and safe responses. This dual-targeting approach exposes weaknesses in both the core LLM and the RM simultaneously. Using the vulnerabilities gained, ARES implements a two-stage repair process: first fine-tuning the RM to better detect harmful content, then leveraging the improved RM to optimize the core model. Experiments across multiple adversarial safety benchmarks demonstrate that ARES substantially enhances safety robustness while preserving model capabilities, establishing a new paradigm for comprehensive RLHF safety alignment.
title ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
topic Artificial Intelligence
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2604.18789