CRISP: Complex Reasoning with Interpretable Step-based Plans

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vetzler, Matan, Lazar, Koren, Uziel, Guy, Hirsch, Eran, Anaby-Tavor, Ateret, Choshen, Leshem
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909683910967296
author Vetzler, Matan
Lazar, Koren
Uziel, Guy
Hirsch, Eran
Anaby-Tavor, Ateret
Choshen, Leshem
author_facet Vetzler, Matan
Lazar, Koren
Uziel, Guy
Hirsch, Eran
Anaby-Tavor, Ateret
Choshen, Leshem
contents Recent advancements in large language models (LLMs) underscore the need for stronger reasoning capabilities to solve complex problems effectively. While Chain-of-Thought (CoT) reasoning has been a step forward, it remains insufficient for many domains. A promising alternative is explicit high-level plan generation, but existing approaches largely assume that LLMs can produce effective plans through few-shot prompting alone, without additional training. In this work, we challenge this assumption and introduce CRISP (Complex Reasoning with Interpretable Step-based Plans), a multi-domain dataset of high-level plans for mathematical reasoning and code generation. The plans in CRISP are automatically generated and rigorously validated--both intrinsically, using an LLM as a judge, and extrinsically, by evaluating their impact on downstream task performance. We demonstrate that fine-tuning a small model on CRISP enables it to generate higher-quality plans than much larger models using few-shot prompting, while significantly outperforming Chain-of-Thought reasoning. Furthermore, our out-of-domain evaluation reveals that fine-tuning on one domain improves plan generation in the other, highlighting the generalizability of learned planning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08037
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CRISP: Complex Reasoning with Interpretable Step-based Plans
Vetzler, Matan
Lazar, Koren
Uziel, Guy
Hirsch, Eran
Anaby-Tavor, Ateret
Choshen, Leshem
Computation and Language
Artificial Intelligence
Recent advancements in large language models (LLMs) underscore the need for stronger reasoning capabilities to solve complex problems effectively. While Chain-of-Thought (CoT) reasoning has been a step forward, it remains insufficient for many domains. A promising alternative is explicit high-level plan generation, but existing approaches largely assume that LLMs can produce effective plans through few-shot prompting alone, without additional training. In this work, we challenge this assumption and introduce CRISP (Complex Reasoning with Interpretable Step-based Plans), a multi-domain dataset of high-level plans for mathematical reasoning and code generation. The plans in CRISP are automatically generated and rigorously validated--both intrinsically, using an LLM as a judge, and extrinsically, by evaluating their impact on downstream task performance. We demonstrate that fine-tuning a small model on CRISP enables it to generate higher-quality plans than much larger models using few-shot prompting, while significantly outperforming Chain-of-Thought reasoning. Furthermore, our out-of-domain evaluation reveals that fine-tuning on one domain improves plan generation in the other, highlighting the generalizability of learned planning capabilities.
title CRISP: Complex Reasoning with Interpretable Step-based Plans
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.08037