QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dineen, Jacob, RRV, Aswin, Liu, Qin, Xu, Zhikun, Ye, Xiao, Shen, Ming, Li, Zhaonan, Lu, Shijie, Baral, Chitta, Chen, Muhao, Zhou, Ben
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912747338334208
author Dineen, Jacob
RRV, Aswin
Liu, Qin
Xu, Zhikun
Ye, Xiao
Shen, Ming
Li, Zhaonan
Lu, Shijie
Baral, Chitta
Chen, Muhao
Zhou, Ben
author_facet Dineen, Jacob
RRV, Aswin
Liu, Qin
Xu, Zhikun
Ye, Xiao
Shen, Ming
Li, Zhaonan
Lu, Shijie
Baral, Chitta
Chen, Muhao
Zhou, Ben
contents Alignment of large language models (LLMs) with principles like helpfulness, honesty, and harmlessness typically relies on scalar rewards that obscure which objectives drive the training signal. We introduce QA-LIGN, which decomposes monolithic rewards into interpretable principle-specific evaluations through structured natural language programs. Models learn through a draft, critique, and revise pipeline, where symbolic evaluation against the rubrics provides transparent feedback for both initial and revised responses during GRPO training. Applied to uncensored Llama-3.1-8B-Instruct, QA-LIGN reduces attack success rates by up to 68.7% while maintaining a 0.67% false refusal rate, achieving Pareto optimal safety-helpfulness performance and outperforming both DPO and GRPO with state-of-the-art reward models given equivalent training. These results demonstrate that making reward signals interpretable and modular improves alignment effectiveness, suggesting transparency enhances LLM safety.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08123
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
Dineen, Jacob
RRV, Aswin
Liu, Qin
Xu, Zhikun
Ye, Xiao
Shen, Ming
Li, Zhaonan
Lu, Shijie
Baral, Chitta
Chen, Muhao
Zhou, Ben
Computation and Language
Alignment of large language models (LLMs) with principles like helpfulness, honesty, and harmlessness typically relies on scalar rewards that obscure which objectives drive the training signal. We introduce QA-LIGN, which decomposes monolithic rewards into interpretable principle-specific evaluations through structured natural language programs. Models learn through a draft, critique, and revise pipeline, where symbolic evaluation against the rubrics provides transparent feedback for both initial and revised responses during GRPO training. Applied to uncensored Llama-3.1-8B-Instruct, QA-LIGN reduces attack success rates by up to 68.7% while maintaining a 0.67% false refusal rate, achieving Pareto optimal safety-helpfulness performance and outperforming both DPO and GRPO with state-of-the-art reward models given equivalent training. These results demonstrate that making reward signals interpretable and modular improves alignment effectiveness, suggesting transparency enhances LLM safety.
title QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
topic Computation and Language
url https://arxiv.org/abs/2506.08123