Checklists Are Better Than Reward Models For Aligning Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Viswanathan, Vijay, Sun, Yanchao, Ma, Shuang, Kong, Xiang, Cao, Meng, Neubig, Graham, Wu, Tongshuang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917115030667264
author Viswanathan, Vijay
Sun, Yanchao
Ma, Shuang
Kong, Xiang
Cao, Meng
Neubig, Graham
Wu, Tongshuang
author_facet Viswanathan, Vijay
Sun, Yanchao
Ma, Shuang
Kong, Xiang
Cao, Meng
Neubig, Graham
Wu, Tongshuang
contents Language models must be adapted to understand and follow user instructions. Reinforcement learning is widely used to facilitate this -- typically using fixed criteria such as "helpfulness" and "harmfulness". In our work, we instead propose using flexible, instruction-specific criteria as a means of broadening the impact that reinforcement learning can have in eliciting instruction following. We propose "Reinforcement Learning from Checklist Feedback" (RLCF). From instructions, we extract checklists and evaluate how well responses satisfy each item - using both AI judges and specialized verifier programs - then combine these scores to compute rewards for RL. We compare RLCF with other alignment methods applied to a strong instruction following model (Qwen2.5-7B-Instruct) on five widely-studied benchmarks -- RLCF is the only method to improve performance on every benchmark, including a 4-point boost in hard satisfaction rate on FollowBench, a 6-point increase on InFoBench, and a 3-point rise in win rate on Arena-Hard. These results establish checklist feedback as a key tool for improving language models' support of queries that express a multitude of needs.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18624
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Checklists Are Better Than Reward Models For Aligning Language Models
Viswanathan, Vijay
Sun, Yanchao
Ma, Shuang
Kong, Xiang
Cao, Meng
Neubig, Graham
Wu, Tongshuang
Computation and Language
Language models must be adapted to understand and follow user instructions. Reinforcement learning is widely used to facilitate this -- typically using fixed criteria such as "helpfulness" and "harmfulness". In our work, we instead propose using flexible, instruction-specific criteria as a means of broadening the impact that reinforcement learning can have in eliciting instruction following. We propose "Reinforcement Learning from Checklist Feedback" (RLCF). From instructions, we extract checklists and evaluate how well responses satisfy each item - using both AI judges and specialized verifier programs - then combine these scores to compute rewards for RL. We compare RLCF with other alignment methods applied to a strong instruction following model (Qwen2.5-7B-Instruct) on five widely-studied benchmarks -- RLCF is the only method to improve performance on every benchmark, including a 4-point boost in hard satisfaction rate on FollowBench, a 6-point increase on InFoBench, and a 3-point rise in win rate on Arena-Hard. These results establish checklist feedback as a key tool for improving language models' support of queries that express a multitude of needs.
title Checklists Are Better Than Reward Models For Aligning Language Models
topic Computation and Language
url https://arxiv.org/abs/2507.18624