Teaching Large Language Models to Reason with Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Havrilla, Alex, Du, Yuqing, Raparthy, Sharath Chandra, Nalmpantis, Christoforos, Dwivedi-Yu, Jane, Zhuravinskyi, Maksym, Hambro, Eric, Sukhbaatar, Sainbayar, Raileanu, Roberta |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
by: Kirk, Robert, et al.
Published: (2023)
by: Kirk, Robert, et al.
Published: (2023)
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
by: Samvelyan, Mikayel, et al.
Published: (2024)
by: Samvelyan, Mikayel, et al.
Published: (2024)
Training Large Language Models to Reason in a Continuous Latent Space
by: Hao, Shibo, et al.
Published: (2024)
by: Hao, Shibo, et al.
Published: (2024)
Self-Challenging Language Model Agents
by: Zhou, Yifei, et al.
Published: (2025)
by: Zhou, Yifei, et al.
Published: (2025)
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
by: Su, DiJia, et al.
Published: (2024)
by: Su, DiJia, et al.
Published: (2024)
Multi-Token Attention
by: Golovneva, Olga, et al.
Published: (2025)
by: Golovneva, Olga, et al.
Published: (2025)
Contextual Position Encoding: Learning to Count What's Important
by: Golovneva, Olga, et al.
Published: (2024)
by: Golovneva, Olga, et al.
Published: (2024)
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
by: Xu, Jing, et al.
Published: (2023)
by: Xu, Jing, et al.
Published: (2023)
IGDA: Interactive Graph Discovery through Large Language Model Agents
by: Havrilla, Alex, et al.
Published: (2025)
by: Havrilla, Alex, et al.
Published: (2025)
Iterative Reasoning Preference Optimization
by: Pang, Richard Yuanzhe, et al.
Published: (2024)
by: Pang, Richard Yuanzhe, et al.
Published: (2024)
TOOLVERIFIER: Generalization to New Tools via Self-Verification
by: Mekala, Dheeraj, et al.
Published: (2024)
by: Mekala, Dheeraj, et al.
Published: (2024)
Reverse Training to Nurse the Reversal Curse
by: Golovneva, Olga, et al.
Published: (2024)
by: Golovneva, Olga, et al.
Published: (2024)
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
Understanding the Effect of Noise in LLM Training Data with Algorithmic Chains of Thought
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
StepWiser: Stepwise Generative Judges for Wiser Reasoning
by: Xiong, Wei, et al.
Published: (2025)
by: Xiong, Wei, et al.
Published: (2025)
The Generalization Gap in Offline Reinforcement Learning
by: Mediratta, Ishita, et al.
Published: (2023)
by: Mediratta, Ishita, et al.
Published: (2023)
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
by: Zhou, Yifei, et al.
Published: (2025)
by: Zhou, Yifei, et al.
Published: (2025)
Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya
by: Sathish, Sharath
Published: (2026)
by: Sathish, Sharath
Published: (2026)
Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
by: Lupidi, Alisia, et al.
Published: (2024)
by: Lupidi, Alisia, et al.
Published: (2024)
SPARQ: Synthetic Problem Generation for Reasoning via Quality-Diversity Algorithms
by: Havrilla, Alex, et al.
Published: (2025)
by: Havrilla, Alex, et al.
Published: (2025)
Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
Self-Rewarding Language Models
by: Yuan, Weizhe, et al.
Published: (2024)
by: Yuan, Weizhe, et al.
Published: (2024)
LLM-First Search: Self-Guided Exploration of the Solution Space
by: Herr, Nathan, et al.
Published: (2025)
by: Herr, Nathan, et al.
Published: (2025)
FairPair: A Robust Evaluation of Biases in Language Models through Paired Perturbations
by: Dwivedi-Yu, Jane, et al.
Published: (2024)
by: Dwivedi-Yu, Jane, et al.
Published: (2024)
Thinking LLMs: General Instruction Following with Thought Generation
by: Wu, Tianhao, et al.
Published: (2024)
by: Wu, Tianhao, et al.
Published: (2024)
Khinchin-type inequalities via Hadamard's factorisation
by: Havrilla, Alex, et al.
Published: (2021)
by: Havrilla, Alex, et al.
Published: (2021)
Following Length Constraints in Instructions
by: Yuan, Weizhe, et al.
Published: (2024)
by: Yuan, Weizhe, et al.
Published: (2024)
Adaptive Decoding via Latent Preference Optimization
by: Dhuliawala, Shehzaad, et al.
Published: (2024)
by: Dhuliawala, Shehzaad, et al.
Published: (2024)
R.I.P.: Better Models by Survival of the Fittest Prompts
by: Yu, Ping, et al.
Published: (2025)
by: Yu, Ping, et al.
Published: (2025)
Diverse Preference Optimization
by: Lanchantin, Jack, et al.
Published: (2025)
by: Lanchantin, Jack, et al.
Published: (2025)
Are Large Language Models Strategic Decision Makers? A Study of Performance and Bias in Two-Player Non-Zero-Sum Games
by: Herr, Nathan, et al.
Published: (2024)
by: Herr, Nathan, et al.
Published: (2024)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
by: Wu, Tianhao, et al.
Published: (2024)
by: Wu, Tianhao, et al.
Published: (2024)
Epistemic Dissonance and Modal Boundaries
by: Raileanu, Dragos
Published: (2025)
by: Raileanu, Dragos
Published: (2025)
Know When To Stop: A Study of Semantic Drift in Text Generation
by: Spataru, Ava, et al.
Published: (2024)
by: Spataru, Ava, et al.
Published: (2024)
SPICE: Self-Play In Corpus Environments Improves Reasoning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Self-Consistency Preference Optimization
by: Prasad, Archiki, et al.
Published: (2024)
by: Prasad, Archiki, et al.
Published: (2024)
A Survey on Hardware Accelerators for Large Language Models
by: Kachris, Christoforos
Published: (2024)
by: Kachris, Christoforos
Published: (2024)
Forgetting as a Feature: Cognitive Alignment of Large Language Models
by: Christoforos, Alexandros
Published: (2025)
by: Christoforos, Alexandros
Published: (2025)
Bridging Offline and Online Reinforcement Learning for LLMs
by: Lanchantin, Jack, et al.
Published: (2025)
by: Lanchantin, Jack, et al.
Published: (2025)
Similar Items
-
GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
by: Havrilla, Alex, et al.
Published: (2024) -
Understanding the Effects of RLHF on LLM Generalisation and Diversity
by: Kirk, Robert, et al.
Published: (2023) -
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
by: Samvelyan, Mikayel, et al.
Published: (2024) -
Training Large Language Models to Reason in a Continuous Latent Space
by: Hao, Shibo, et al.
Published: (2024) -
Self-Challenging Language Model Agents
by: Zhou, Yifei, et al.
Published: (2025)