Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Parashar, Shubham, Olson, Blake, Khurana, Sambhav, Li, Eric, Ling, Hongyi, Caverlee, James, Ji, Shuiwang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913695759597568
author Parashar, Shubham
Olson, Blake
Khurana, Sambhav
Li, Eric
Ling, Hongyi
Caverlee, James
Ji, Shuiwang
author_facet Parashar, Shubham
Olson, Blake
Khurana, Sambhav
Li, Eric
Ling, Hongyi
Caverlee, James
Ji, Shuiwang
contents We examine the reasoning and planning capabilities of large language models (LLMs) in solving complex tasks. Recent advances in inference-time techniques demonstrate the potential to enhance LLM reasoning without additional training by exploring intermediate steps during inference. Notably, OpenAI's o1 model shows promising performance through its novel use of multi-step reasoning and verification. Here, we explore how scaling inference-time techniques can improve reasoning and planning, focusing on understanding the tradeoff between computational cost and performance. To this end, we construct a comprehensive benchmark, known as Sys2Bench, and perform extensive experiments evaluating existing inference-time techniques on eleven diverse tasks across five categories, including arithmetic reasoning, logical reasoning, common sense reasoning, algorithmic reasoning, and planning. Our findings indicate that simply scaling inference-time computation has limitations, as no single inference-time technique consistently performs well across all reasoning and planning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12521
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights
Parashar, Shubham
Olson, Blake
Khurana, Sambhav
Li, Eric
Ling, Hongyi
Caverlee, James
Ji, Shuiwang
Artificial Intelligence
Machine Learning
We examine the reasoning and planning capabilities of large language models (LLMs) in solving complex tasks. Recent advances in inference-time techniques demonstrate the potential to enhance LLM reasoning without additional training by exploring intermediate steps during inference. Notably, OpenAI's o1 model shows promising performance through its novel use of multi-step reasoning and verification. Here, we explore how scaling inference-time techniques can improve reasoning and planning, focusing on understanding the tradeoff between computational cost and performance. To this end, we construct a comprehensive benchmark, known as Sys2Bench, and perform extensive experiments evaluating existing inference-time techniques on eleven diverse tasks across five categories, including arithmetic reasoning, logical reasoning, common sense reasoning, algorithmic reasoning, and planning. Our findings indicate that simply scaling inference-time computation has limitations, as no single inference-time technique consistently performs well across all reasoning and planning tasks.
title Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.12521