Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Badola, Kartikeya, Simon, Jonathan, Hosseini, Arian, Carthy, Sara Marie Mc, Munkhdalai, Tsendsuren, Goyal, Abhimanyu, Kočiský, Tomáš, Upadhyay, Shyam, Fatemi, Bahare, Kazemi, Mehran
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909750584672256
author Badola, Kartikeya
Simon, Jonathan
Hosseini, Arian
Carthy, Sara Marie Mc
Munkhdalai, Tsendsuren
Goyal, Abhimanyu
Kočiský, Tomáš
Upadhyay, Shyam
Fatemi, Bahare
Kazemi, Mehran
author_facet Badola, Kartikeya
Simon, Jonathan
Hosseini, Arian
Carthy, Sara Marie Mc
Munkhdalai, Tsendsuren
Goyal, Abhimanyu
Kočiský, Tomáš
Upadhyay, Shyam
Fatemi, Bahare
Kazemi, Mehran
contents Large language models (LLMs) excel at solving problems with clear and complete statements, but often struggle with nuanced environments or interactive tasks which are common in most real-world scenarios. This highlights the critical need for developing LLMs that can effectively engage in logically consistent multi-turn dialogue, seek information and reason with incomplete data. To this end, we introduce a novel benchmark comprising a suite of multi-turn tasks each designed to test specific reasoning, interactive dialogue, and information-seeking abilities. These tasks have deterministic scoring mechanisms, thus eliminating the need for human intervention. Evaluating frontier models on our benchmark reveals significant headroom. Our analysis shows that most errors emerge from poor instruction following, reasoning failures, and poor planning. This benchmark provides valuable insights into the strengths and weaknesses of current LLMs in handling complex, interactive scenarios and offers a robust platform for future research aimed at improving these critical capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10142
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
Badola, Kartikeya
Simon, Jonathan
Hosseini, Arian
Carthy, Sara Marie Mc
Munkhdalai, Tsendsuren
Goyal, Abhimanyu
Kočiský, Tomáš
Upadhyay, Shyam
Fatemi, Bahare
Kazemi, Mehran
Computation and Language
Large language models (LLMs) excel at solving problems with clear and complete statements, but often struggle with nuanced environments or interactive tasks which are common in most real-world scenarios. This highlights the critical need for developing LLMs that can effectively engage in logically consistent multi-turn dialogue, seek information and reason with incomplete data. To this end, we introduce a novel benchmark comprising a suite of multi-turn tasks each designed to test specific reasoning, interactive dialogue, and information-seeking abilities. These tasks have deterministic scoring mechanisms, thus eliminating the need for human intervention. Evaluating frontier models on our benchmark reveals significant headroom. Our analysis shows that most errors emerge from poor instruction following, reasoning failures, and poor planning. This benchmark provides valuable insights into the strengths and weaknesses of current LLMs in handling complex, interactive scenarios and offers a robust platform for future research aimed at improving these critical capabilities.
title Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
topic Computation and Language
url https://arxiv.org/abs/2508.10142