Don't Walk the Line: Boundary Guidance for Filtered Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ball, Sarah, Haupt, Andreas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911444195344384
author Ball, Sarah
Haupt, Andreas
author_facet Ball, Sarah
Haupt, Andreas
contents Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tune the generator to reduce the probability of being filtered, but this can be suboptimal: it often pushes the model toward producing samples near the classifier's decision boundary, increasing both false positives and false negatives. We propose Boundary Guidance, a reinforcement learning fine-tuning method that explicitly steers generation away from the classifier's margin. On a benchmark of jailbreak, ambiguous, and longcontext prompts, Boundary Guidance improves both the safety and the utility of outputs, as judged by LLM-as-a-Judge evaluations. Comprehensive ablations across model scales and reward designs demonstrate the robustness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11834
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Don't Walk the Line: Boundary Guidance for Filtered Generation
Ball, Sarah
Haupt, Andreas
Machine Learning
Computation and Language
Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tune the generator to reduce the probability of being filtered, but this can be suboptimal: it often pushes the model toward producing samples near the classifier's decision boundary, increasing both false positives and false negatives. We propose Boundary Guidance, a reinforcement learning fine-tuning method that explicitly steers generation away from the classifier's margin. On a benchmark of jailbreak, ambiguous, and longcontext prompts, Boundary Guidance improves both the safety and the utility of outputs, as judged by LLM-as-a-Judge evaluations. Comprehensive ablations across model scales and reward designs demonstrate the robustness of our approach.
title Don't Walk the Line: Boundary Guidance for Filtered Generation
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2510.11834