Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Zeyu, Dai, Xiangxiang, Han, Ziyi, Liu, Xutong, Lui, John C. S.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917348469899264
author Zhang, Zeyu
Dai, Xiangxiang
Han, Ziyi
Liu, Xutong
Lui, John C. S.
author_facet Zhang, Zeyu
Dai, Xiangxiang
Han, Ziyi
Liu, Xutong
Lui, John C. S.
contents Large language models (LLMs) are typically governed by post-training alignment (e.g., RLHF or DPO), which yields a largely static policy during deployment and inference. However, real-world safety is a full-lifecycle problem: static defenses degrade against evolving jailbreak behaviors, and fixed weights cannot adapt to pluralistic, time-varying safety norms. This motivates inference-time governance that steers behavior without costly retraining. To address this, we introduce the Consensus Clustering LinUCB Bandit (CCLUB), a unified framework for adaptive social alignment via system-prompt routing. CCLUB employs a conservative consensus clustering mechanism: it pools data only within the intersection of utility and safety similarity graphs, effectively preventing unsafe generalization across semantically proximal but risk-divergent contexts. Our theoretical analysis yields a sublinear regret guarantee, demonstrating near-optimal performance of CCLUB. Extensive experiments validate that CCLUB outperforms strong baselines, achieving a 10.98% improvement in cumulative reward and a 14.42% reduction in the average suboptimality gap.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15647
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
Zhang, Zeyu
Dai, Xiangxiang
Han, Ziyi
Liu, Xutong
Lui, John C. S.
Machine Learning
Artificial Intelligence
Large language models (LLMs) are typically governed by post-training alignment (e.g., RLHF or DPO), which yields a largely static policy during deployment and inference. However, real-world safety is a full-lifecycle problem: static defenses degrade against evolving jailbreak behaviors, and fixed weights cannot adapt to pluralistic, time-varying safety norms. This motivates inference-time governance that steers behavior without costly retraining. To address this, we introduce the Consensus Clustering LinUCB Bandit (CCLUB), a unified framework for adaptive social alignment via system-prompt routing. CCLUB employs a conservative consensus clustering mechanism: it pools data only within the intersection of utility and safety similarity graphs, effectively preventing unsafe generalization across semantically proximal but risk-divergent contexts. Our theoretical analysis yields a sublinear regret guarantee, demonstrating near-optimal performance of CCLUB. Extensive experiments validate that CCLUB outperforms strong baselines, achieving a 10.98% improvement in cumulative reward and a 14.42% reduction in the average suboptimality gap.
title Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.15647