Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ren, Qingyu, He, Qianyu, Chang, Powei, Zeng, Jie, Sun, Zeye, Yu, Fei, Liang, Jiaqing, Xiao, Yanghua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910126984658944
author Ren, Qingyu
He, Qianyu
Chang, Powei
Zeng, Jie
Sun, Zeye
Yu, Fei
Liang, Jiaqing
Xiao, Yanghua
author_facet Ren, Qingyu
He, Qianyu
Chang, Powei
Zeng, Jie
Sun, Zeye
Yu, Fei
Liang, Jiaqing
Xiao, Yanghua
contents Language models often struggle to follow multi-constraint instructions that are crucial for real-world applications. Existing reinforcement learning (RL) approaches suffer from dependency on external supervision and sparse reward signals from multi-constraint tasks. We propose a label-free self-supervised RL framework that eliminates dependency on external supervision by deriving reward signals directly from instructions and generating pseudo-labels for reward model training. Our approach introduces constraint decomposition strategies and efficient constraint-wise binary classification to address sparse reward challenges while maintaining computational efficiency. Experiments show that our approach generalizes well, achieving strong improvements across 3 in-domain and 5 out-of-domain datasets, including challenging agentic and multi-turn instruction following. The data and code are publicly available at https://github.com/Rainier-rq/verl-if
format Preprint
id arxiv_https___arxiv_org_abs_2510_14420
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
Ren, Qingyu
He, Qianyu
Chang, Powei
Zeng, Jie
Sun, Zeye
Yu, Fei
Liang, Jiaqing
Xiao, Yanghua
Computation and Language
Artificial Intelligence
Language models often struggle to follow multi-constraint instructions that are crucial for real-world applications. Existing reinforcement learning (RL) approaches suffer from dependency on external supervision and sparse reward signals from multi-constraint tasks. We propose a label-free self-supervised RL framework that eliminates dependency on external supervision by deriving reward signals directly from instructions and generating pseudo-labels for reward model training. Our approach introduces constraint decomposition strategies and efficient constraint-wise binary classification to address sparse reward challenges while maintaining computational efficiency. Experiments show that our approach generalizes well, achieving strong improvements across 3 in-domain and 5 out-of-domain datasets, including challenging agentic and multi-turn instruction following. The data and code are publicly available at https://github.com/Rainier-rq/verl-if
title Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.14420