Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Yonghui, Tao, Wenjian, Liu, Jilong, Zhu, Xingyu, Fang, Junfeng, Huang, Weibiao, Wu, Le, Hong, Richang, Chua, Tat-Sent
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916034533916672
author Yang, Yonghui
Tao, Wenjian
Liu, Jilong
Zhu, Xingyu
Fang, Junfeng
Huang, Weibiao
Wu, Le
Hong, Richang
Chua, Tat-Sent
author_facet Yang, Yonghui
Tao, Wenjian
Liu, Jilong
Zhu, Xingyu
Fang, Junfeng
Huang, Weibiao
Wu, Le
Hong, Richang
Chua, Tat-Sent
contents Safety alignment of large language models remains brittle under domain shift and noisy preference supervision. Most existing robust alignment methods focus on uncertainty in alignment data, while overlooking optimization-induced fragility in preference-based objectives. In this work, we revisit robustness for LLM safety alignment from an optimization geometry perspective, and argue that robustness failures cannot be addressed by data-centric methods alone. We propose \textit{ShaPO}, a geometry-aware preference optimization framework that enforces worst-case alignment objectives via selective geometry control over alignment-critical parameter subspace. By avoiding uniform geometry constraints, ShaPO mitigates the over-regularization that can harm robustness under distribution shift. We instantiate ShaPO at two levels: token-level ShaPO stabilizes likelihood-based surrogate optimization, while reward-level ShaPO enforces reward-consistent optimization under noisy supervision. Across diverse safety benchmarks and noisy preference settings, ShaPO consistently improves safety robustness over popular preference optimization methods. Moreover, ShaPO composes cleanly with data-robust objectives, yielding additional gains and empirically supporting the proposed optimization-geometry perspective. The code is available at https://github.com/liujilong0116/ShaPO.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07340
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
Yang, Yonghui
Tao, Wenjian
Liu, Jilong
Zhu, Xingyu
Fang, Junfeng
Huang, Weibiao
Wu, Le
Hong, Richang
Chua, Tat-Sent
Machine Learning
Safety alignment of large language models remains brittle under domain shift and noisy preference supervision. Most existing robust alignment methods focus on uncertainty in alignment data, while overlooking optimization-induced fragility in preference-based objectives. In this work, we revisit robustness for LLM safety alignment from an optimization geometry perspective, and argue that robustness failures cannot be addressed by data-centric methods alone. We propose \textit{ShaPO}, a geometry-aware preference optimization framework that enforces worst-case alignment objectives via selective geometry control over alignment-critical parameter subspace. By avoiding uniform geometry constraints, ShaPO mitigates the over-regularization that can harm robustness under distribution shift. We instantiate ShaPO at two levels: token-level ShaPO stabilizes likelihood-based surrogate optimization, while reward-level ShaPO enforces reward-consistent optimization under noisy supervision. Across diverse safety benchmarks and noisy preference settings, ShaPO consistently improves safety robustness over popular preference optimization methods. Moreover, ShaPO composes cleanly with data-robust objectives, yielding additional gains and empirically supporting the proposed optimization-geometry perspective. The code is available at https://github.com/liujilong0116/ShaPO.
title Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
topic Machine Learning
url https://arxiv.org/abs/2602.07340