QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912747338334208 |
|---|---|
| author | Dineen, Jacob RRV, Aswin Liu, Qin Xu, Zhikun Ye, Xiao Shen, Ming Li, Zhaonan Lu, Shijie Baral, Chitta Chen, Muhao Zhou, Ben |
| author_facet | Dineen, Jacob RRV, Aswin Liu, Qin Xu, Zhikun Ye, Xiao Shen, Ming Li, Zhaonan Lu, Shijie Baral, Chitta Chen, Muhao Zhou, Ben |
| contents | Alignment of large language models (LLMs) with principles like helpfulness, honesty, and harmlessness typically relies on scalar rewards that obscure which objectives drive the training signal. We introduce QA-LIGN, which decomposes monolithic rewards into interpretable principle-specific evaluations through structured natural language programs. Models learn through a draft, critique, and revise pipeline, where symbolic evaluation against the rubrics provides transparent feedback for both initial and revised responses during GRPO training. Applied to uncensored Llama-3.1-8B-Instruct, QA-LIGN reduces attack success rates by up to 68.7% while maintaining a 0.67% false refusal rate, achieving Pareto optimal safety-helpfulness performance and outperforming both DPO and GRPO with state-of-the-art reward models given equivalent training. These results demonstrate that making reward signals interpretable and modular improves alignment effectiveness, suggesting transparency enhances LLM safety. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_08123 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA Dineen, Jacob RRV, Aswin Liu, Qin Xu, Zhikun Ye, Xiao Shen, Ming Li, Zhaonan Lu, Shijie Baral, Chitta Chen, Muhao Zhou, Ben Computation and Language Alignment of large language models (LLMs) with principles like helpfulness, honesty, and harmlessness typically relies on scalar rewards that obscure which objectives drive the training signal. We introduce QA-LIGN, which decomposes monolithic rewards into interpretable principle-specific evaluations through structured natural language programs. Models learn through a draft, critique, and revise pipeline, where symbolic evaluation against the rubrics provides transparent feedback for both initial and revised responses during GRPO training. Applied to uncensored Llama-3.1-8B-Instruct, QA-LIGN reduces attack success rates by up to 68.7% while maintaining a 0.67% false refusal rate, achieving Pareto optimal safety-helpfulness performance and outperforming both DPO and GRPO with state-of-the-art reward models given equivalent training. These results demonstrate that making reward signals interpretable and modular improves alignment effectiveness, suggesting transparency enhances LLM safety. |
| title | QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2506.08123 |