Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Haozheng, Wang, Yimin, Yu, Jiahao, Wang, Binghui, Chen, Yan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910235640201216
author Luo, Haozheng
Wang, Yimin
Yu, Jiahao
Wang, Binghui
Chen, Yan
author_facet Luo, Haozheng
Wang, Yimin
Yu, Jiahao
Wang, Binghui
Chen, Yan
contents We propose CRAFT, a red-teaming alignment framework that leverages model reasoning capabilities and hidden representations to improve robustness against jailbreak attacks. Unlike prior defenses that operate primarily at the output level, CRAFT aligns large reasoning models to generate safety-aware reasoning traces by explicitly optimizing objectives defined over the hidden state space. Methodologically, CRAFT integrates contrastive representation learning with reinforcement learning to separate safe and unsafe reasoning trajectories, yielding a latent-space geometry that supports robust, reasoning-level safety alignment. Theoretically, we show that incorporating latent-textual consistency into GRPO eliminates superficially aligned policies by ruling them out as local optima. Empirically, we evaluate CRAFT on multiple safety benchmarks using two strong reasoning models, Qwen3-4B-Thinking and R1-Distill-Llama-8B, where it consistently outperforms state-of-the-art defenses such as IPO and SafeKey. Notably, CRAFT delivers an average 79.0% improvement in reasoning safety and 87.7% improvement in final-response safety over the base models, demonstrating the effectiveness of hidden-space reasoning alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17305
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
Luo, Haozheng
Wang, Yimin
Yu, Jiahao
Wang, Binghui
Chen, Yan
Artificial Intelligence
Computation and Language
Machine Learning
We propose CRAFT, a red-teaming alignment framework that leverages model reasoning capabilities and hidden representations to improve robustness against jailbreak attacks. Unlike prior defenses that operate primarily at the output level, CRAFT aligns large reasoning models to generate safety-aware reasoning traces by explicitly optimizing objectives defined over the hidden state space. Methodologically, CRAFT integrates contrastive representation learning with reinforcement learning to separate safe and unsafe reasoning trajectories, yielding a latent-space geometry that supports robust, reasoning-level safety alignment. Theoretically, we show that incorporating latent-textual consistency into GRPO eliminates superficially aligned policies by ruling them out as local optima. Empirically, we evaluate CRAFT on multiple safety benchmarks using two strong reasoning models, Qwen3-4B-Thinking and R1-Distill-Llama-8B, where it consistently outperforms state-of-the-art defenses such as IPO and SafeKey. Notably, CRAFT delivers an average 79.0% improvement in reasoning safety and 87.7% improvement in final-response safety over the base models, demonstrating the effectiveness of hidden-space reasoning alignment.
title Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2603.17305