Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Qingyu, He, Qianyu, Chang, Powei, Zeng, Jie, Sun, Zeye, Yu, Fei, Liang, Jiaqing, Xiao, Yanghua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following
by: Ren, Qingyu, et al.
Published: (2025)
by: Ren, Qingyu, et al.
Published: (2025)
SEIF: Self-Evolving Reinforcement Learning for Instruction Following
by: Ren, Qingyu, et al.
Published: (2026)
by: Ren, Qingyu, et al.
Published: (2026)
Order Matters: Investigate the Position Bias in Multi-constraint Instruction Following
by: Zeng, Jie, et al.
Published: (2025)
by: Zeng, Jie, et al.
Published: (2025)
Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models
by: Ren, Qingyu, et al.
Published: (2025)
by: Ren, Qingyu, et al.
Published: (2025)
LsrIF: Enhancing Logic-Structured Instruction Following of Large Language Models
by: Ren, Qingyu, et al.
Published: (2026)
by: Ren, Qingyu, et al.
Published: (2026)
From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models
by: He, Qianyu, et al.
Published: (2024)
by: He, Qianyu, et al.
Published: (2024)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
by: He, Qianxi, et al.
Published: (2025)
by: He, Qianxi, et al.
Published: (2025)
Small Language Model Can Self-correct
by: Han, Haixia, et al.
Published: (2024)
by: Han, Haixia, et al.
Published: (2024)
Laying the Foundation First? Investigating the Generalization from Atomic Skills to Complex Reasoning Tasks
by: Huang, Yuncheng, et al.
Published: (2024)
by: Huang, Yuncheng, et al.
Published: (2024)
Is There a One-Model-Fits-All Approach to Information Extraction? Revisiting Task Definition Biases
by: Huang, Wenhao, et al.
Published: (2024)
by: Huang, Wenhao, et al.
Published: (2024)
ChemAmp: Amplified Chemistry Tools via Composable Agents
by: Li, Zhucong, et al.
Published: (2025)
by: Li, Zhucong, et al.
Published: (2025)
Can Large Language Models Understand Real-World Complex Instructions?
by: He, Qianyu, et al.
Published: (2023)
by: He, Qianyu, et al.
Published: (2023)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
by: Pan, Tianjun, et al.
Published: (2026)
by: Pan, Tianjun, et al.
Published: (2026)
A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models
by: Han, Jinyi, et al.
Published: (2025)
by: Han, Jinyi, et al.
Published: (2025)
Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension Perception
by: Huang, Yuncheng, et al.
Published: (2023)
by: Huang, Yuncheng, et al.
Published: (2023)
Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
by: Zhang, Yikai, et al.
Published: (2024)
by: Zhang, Yikai, et al.
Published: (2024)
Adaptive Ordered Information Extraction with Deep Reinforcement Learning
by: Huang, Wenhao, et al.
Published: (2023)
by: Huang, Wenhao, et al.
Published: (2023)
What Makes an Ideal Quote? Recommending "Unexpected yet Rational" Quotations via Novelty
by: Zhang, Bowei, et al.
Published: (2025)
by: Zhang, Bowei, et al.
Published: (2025)
Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding
by: Li, Yanda, et al.
Published: (2024)
by: Li, Yanda, et al.
Published: (2024)
aim_is_all_you_need
by: Katz, Seth
Published: (2025)
by: Katz, Seth
Published: (2025)
Discernment is all you need
by: Fuenmayor, David
Published: (2026)
by: Fuenmayor, David
Published: (2026)
Propaganda is all you need
by: Kronlund-Drouault, Paul
Published: (2024)
by: Kronlund-Drouault, Paul
Published: (2024)
Instruction Following without Instruction Tuning
by: Hewitt, John, et al.
Published: (2024)
by: Hewitt, John, et al.
Published: (2024)
IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
by: Guo, Xu, et al.
Published: (2025)
by: Guo, Xu, et al.
Published: (2025)
Attention is all you need to solve chiral superconductivity
by: Li, Chun-Tse, et al.
Published: (2025)
by: Li, Chun-Tse, et al.
Published: (2025)
Tabular Data: Is Deep Learning all you need?
by: Zabërgja, Guri, et al.
Published: (2024)
by: Zabërgja, Guri, et al.
Published: (2024)
One protein is all you need
by: Bushuiev, Anton, et al.
Published: (2024)
by: Bushuiev, Anton, et al.
Published: (2024)
When Agents Evolve, Institutions Follow
by: Fei, Chao, et al.
Published: (2026)
by: Fei, Chao, et al.
Published: (2026)
QUILL: Quotation Generation Enhancement of Large Language Models
by: Xiao, Jin, et al.
Published: (2024)
by: Xiao, Jin, et al.
Published: (2024)
Is attention all you need to solve the correlated electron problem?
by: Geier, Max, et al.
Published: (2025)
by: Geier, Max, et al.
Published: (2025)
Compositional Instruction Following with Language Models and Reinforcement Learning
by: Cohen, Vanya, et al.
Published: (2025)
by: Cohen, Vanya, et al.
Published: (2025)
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
by: Peng, Hao, et al.
Published: (2025)
by: Peng, Hao, et al.
Published: (2025)
Self-Guided Plan Extraction for Instruction-Following Tasks with Goal-Conditional Reinforcement Learning
by: Volovikova, Zoya, et al.
Published: (2026)
by: Volovikova, Zoya, et al.
Published: (2026)
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
by: Jiang, Zishang, et al.
Published: (2025)
by: Jiang, Zishang, et al.
Published: (2025)
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
by: Zhang, Zhicheng, et al.
Published: (2025)
by: Zhang, Zhicheng, et al.
Published: (2025)
Learning or Self-aligning? Rethinking Instruction Fine-tuning
by: Ren, Mengjie, et al.
Published: (2024)
by: Ren, Mengjie, et al.
Published: (2024)
REL: Working out is all you need
by: Simonds, Toby, et al.
Published: (2024)
by: Simonds, Toby, et al.
Published: (2024)
Filtered Rayleigh-Ritz is all you need
by: Abbott, Ryan, et al.
Published: (2025)
by: Abbott, Ryan, et al.
Published: (2025)
Kolmogorov GAM Networks are all you need!
by: Polson, Sarah, et al.
Published: (2025)
by: Polson, Sarah, et al.
Published: (2025)
Anti-concentration is (almost) all you need
by: Heinrich, Markus, et al.
Published: (2025)
by: Heinrich, Markus, et al.
Published: (2025)
Similar Items
-
Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following
by: Ren, Qingyu, et al.
Published: (2025) -
SEIF: Self-Evolving Reinforcement Learning for Instruction Following
by: Ren, Qingyu, et al.
Published: (2026) -
Order Matters: Investigate the Position Bias in Multi-constraint Instruction Following
by: Zeng, Jie, et al.
Published: (2025) -
Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models
by: Ren, Qingyu, et al.
Published: (2025) -
LsrIF: Enhancing Logic-Structured Instruction Following of Large Language Models
by: Ren, Qingyu, et al.
Published: (2026)