Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Zhenhao, Chang, Wenhan, Chen, Yichuan, Fang, Yuxin, Liu, Junhao, Zhu, Tianqing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913116015558656
author Xu, Zhenhao
Chang, Wenhan
Chen, Yichuan
Fang, Yuxin
Liu, Junhao
Zhu, Tianqing
author_facet Xu, Zhenhao
Chang, Wenhan
Chen, Yichuan
Fang, Yuxin
Liu, Junhao
Zhu, Tianqing
contents Large Reasoning Models (LRMs) improve performance on complex tasks, but they also make safety control harder at deployment time. In black-box settings, defenders cannot modify model weights and must instead intervene at inference time. This setting creates three practical challenges: harmful intent may be hidden by educational or role-play framing, deep safety analysis can introduce non-trivial latency, and long adversarial contexts can dilute the local cues that simpler filters rely on. These challenges can expose an apparent thinking--output gap, where the model appears cautious during reasoning but still produces an unsafe final answer. To address this problem, we propose Safety Context Injection (SCI), an inference-time framework that separates safety assessment from task generation and prepends a structured external risk report as injected safety context for the protected model. The framework is instantiated in two complementary variants: Static Model Filtering (SMF), a lightweight one-pass guard for fast deployment, and Dynamic Agents Filtering (DAF), an agentic-loop-based analyzer that iteratively gathers and synthesizes evidence for ambiguous or long-context attacks. Across AdvBench and GPTFuzz, spanning base and reasoning models under five jailbreak families, both variants reduce attack success rate and toxicity in the evaluated settings. SMF offers an efficient low-latency option, while DAF is more effective when harmful intent is semantically disguised or dispersed across long contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11664
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
Xu, Zhenhao
Chang, Wenhan
Chen, Yichuan
Fang, Yuxin
Liu, Junhao
Zhu, Tianqing
Cryptography and Security
Large Reasoning Models (LRMs) improve performance on complex tasks, but they also make safety control harder at deployment time. In black-box settings, defenders cannot modify model weights and must instead intervene at inference time. This setting creates three practical challenges: harmful intent may be hidden by educational or role-play framing, deep safety analysis can introduce non-trivial latency, and long adversarial contexts can dilute the local cues that simpler filters rely on. These challenges can expose an apparent thinking--output gap, where the model appears cautious during reasoning but still produces an unsafe final answer. To address this problem, we propose Safety Context Injection (SCI), an inference-time framework that separates safety assessment from task generation and prepends a structured external risk report as injected safety context for the protected model. The framework is instantiated in two complementary variants: Static Model Filtering (SMF), a lightweight one-pass guard for fast deployment, and Dynamic Agents Filtering (DAF), an agentic-loop-based analyzer that iteratively gathers and synthesizes evidence for ambiguous or long-context attacks. Across AdvBench and GPTFuzz, spanning base and reasoning models under five jailbreak families, both variants reduce attack success rate and toxicity in the evaluated settings. SMF offers an efficient low-latency option, while DAF is more effective when harmful intent is semantically disguised or dispersed across long contexts.
title Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
topic Cryptography and Security
url https://arxiv.org/abs/2605.11664