SLR: Automated Synthesis for Scalable Logical Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Helff, Lukas, Omar, Ahmad, Friedrich, Felix, Wüst, Antonia, Shindo, Hikaru, Mitchell, Rupert, Woydt, Tim, Schramowski, Patrick, Stammer, Wolfgang, Kersting, Kristian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911668129234944
author Helff, Lukas
Omar, Ahmad
Friedrich, Felix
Wüst, Antonia
Shindo, Hikaru
Mitchell, Rupert
Woydt, Tim
Schramowski, Patrick
Stammer, Wolfgang
Kersting, Kristian
author_facet Helff, Lukas
Omar, Ahmad
Friedrich, Felix
Wüst, Antonia
Shindo, Hikaru
Mitchell, Rupert
Woydt, Tim
Schramowski, Patrick
Stammer, Wolfgang
Kersting, Kristian
contents We introduce SLR, an end-to-end framework for systematic evaluation and training of Large Language Models (LLMs) via Scalable Logical Reasoning. Given a user's task specification, SLR automatically synthesizes (i) an instruction prompt for an inductive reasoning task, (ii) a validation program, executable on model outputs to provide verifiable rewards, and (iii) the latent ground-truth rule. This process is fully automated, scalable, requires no human annotations, and offers precise control over task difficulty. Using SLR, we create SLR-Bench, a benchmark comprising 19k prompts organized into 20 curriculum levels that progressively increase in relational, arithmetic, and recursive complexity. Large-scale evaluation reveals that contemporary LLMs readily produce syntactically valid rules, yet often fail at correct logical inference. Recent reasoning LLMs demonstrate improved performance but incur very high test-time computation, with costs exceeding $300 for just 1,000 prompts. Finally, curriculum learning via SLR doubles Llama-3-8B accuracy on SLR-Bench, achieving parity with Gemini-Flash-Thinking at a fraction of computational cost. Moreover, these reasoning capabilities generalize to a wide range of established benchmarks, underscoring the effectiveness of SLR for downstream reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15787
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SLR: Automated Synthesis for Scalable Logical Reasoning
Helff, Lukas
Omar, Ahmad
Friedrich, Felix
Wüst, Antonia
Shindo, Hikaru
Mitchell, Rupert
Woydt, Tim
Schramowski, Patrick
Stammer, Wolfgang
Kersting, Kristian
Artificial Intelligence
Computation and Language
Machine Learning
We introduce SLR, an end-to-end framework for systematic evaluation and training of Large Language Models (LLMs) via Scalable Logical Reasoning. Given a user's task specification, SLR automatically synthesizes (i) an instruction prompt for an inductive reasoning task, (ii) a validation program, executable on model outputs to provide verifiable rewards, and (iii) the latent ground-truth rule. This process is fully automated, scalable, requires no human annotations, and offers precise control over task difficulty. Using SLR, we create SLR-Bench, a benchmark comprising 19k prompts organized into 20 curriculum levels that progressively increase in relational, arithmetic, and recursive complexity. Large-scale evaluation reveals that contemporary LLMs readily produce syntactically valid rules, yet often fail at correct logical inference. Recent reasoning LLMs demonstrate improved performance but incur very high test-time computation, with costs exceeding $300 for just 1,000 prompts. Finally, curriculum learning via SLR doubles Llama-3-8B accuracy on SLR-Bench, achieving parity with Gemini-Flash-Thinking at a fraction of computational cost. Moreover, these reasoning capabilities generalize to a wide range of established benchmarks, underscoring the effectiveness of SLR for downstream reasoning.
title SLR: Automated Synthesis for Scalable Logical Reasoning
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.15787