Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lacombe, Valentin, Quesnel, Valentin, Sileo, Damien
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909799890812928
author Lacombe, Valentin
Quesnel, Valentin
Sileo, Damien
author_facet Lacombe, Valentin
Quesnel, Valentin
Sileo, Damien
contents We introduce Reasoning Core, a new scalable environment for Reinforcement Learning with Verifiable Rewards (RLVR), designed to advance foundational symbolic reasoning in Large Language Models (LLMs). Unlike existing benchmarks that focus on games or isolated puzzles, Reasoning Core procedurally generates problems across core formal domains, including PDDL planning, first-order logic, context-free grammar parsing, causal reasoning, and system equation solving. The environment is built on key design principles of high-generality problem distributions, verification via external tools, and continuous difficulty control, which together provide a virtually infinite supply of novel training instances. Initial zero-shot evaluations with frontier LLMs confirm the difficulty of Reasoning Core's tasks, positioning it as a promising resource to improve the reasoning capabilities of future models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18083
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning
Lacombe, Valentin
Quesnel, Valentin
Sileo, Damien
Artificial Intelligence
Computation and Language
We introduce Reasoning Core, a new scalable environment for Reinforcement Learning with Verifiable Rewards (RLVR), designed to advance foundational symbolic reasoning in Large Language Models (LLMs). Unlike existing benchmarks that focus on games or isolated puzzles, Reasoning Core procedurally generates problems across core formal domains, including PDDL planning, first-order logic, context-free grammar parsing, causal reasoning, and system equation solving. The environment is built on key design principles of high-generality problem distributions, verification via external tools, and continuous difficulty control, which together provide a virtually infinite supply of novel training instances. Initial zero-shot evaluations with frontier LLMs confirm the difficulty of Reasoning Core's tasks, positioning it as a promising resource to improve the reasoning capabilities of future models.
title Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.18083