Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Badola, Kartikeya, Simon, Jonathan, Hosseini, Arian, Carthy, Sara Marie Mc, Munkhdalai, Tsendsuren, Goyal, Abhimanyu, Kočiský, Tomáš, Upadhyay, Shyam, Fatemi, Bahare, Kazemi, Mehran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
von: Munkhdalai, Tsendsuren, et al.
Veröffentlicht: (2024)
von: Munkhdalai, Tsendsuren, et al.
Veröffentlicht: (2024)
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2025)
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2025)
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
von: Fatemi, Bahare, et al.
Veröffentlicht: (2024)
von: Fatemi, Bahare, et al.
Veröffentlicht: (2024)
Let Your Graph Do the Talking: Encoding Structured Data for LLMs
von: Perozzi, Bryan, et al.
Veröffentlicht: (2024)
von: Perozzi, Bryan, et al.
Veröffentlicht: (2024)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Understanding Transformer Reasoning Capabilities via Graph Algorithms
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
CausalGraph2LLM: Evaluating LLMs for Causal Queries
von: Sheth, Ivaxi, et al.
Veröffentlicht: (2024)
von: Sheth, Ivaxi, et al.
Veröffentlicht: (2024)
Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
von: Hahn, Meera, et al.
Veröffentlicht: (2024)
von: Hahn, Meera, et al.
Veröffentlicht: (2024)
Generative Verifiers: Reward Modeling as Next-Token Prediction
von: Zhang, Lunjun, et al.
Veröffentlicht: (2024)
von: Zhang, Lunjun, et al.
Veröffentlicht: (2024)
ReMI: A Dataset for Reasoning with Multiple Images
von: Kazemi, Mehran, et al.
Veröffentlicht: (2024)
von: Kazemi, Mehran, et al.
Veröffentlicht: (2024)
Hierarchical Recurrent Adapters for Efficient Multi-Task Adaptation of Large Speech Models
von: Munkhdalai, Tsendsuren, et al.
Veröffentlicht: (2024)
von: Munkhdalai, Tsendsuren, et al.
Veröffentlicht: (2024)
What Matters for Model Merging at Scale?
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
CONFEX: Uncertainty-Aware Counterfactual Explanations with Conformal Guarantees
von: Bilkhoo, Aman, et al.
Veröffentlicht: (2025)
von: Bilkhoo, Aman, et al.
Veröffentlicht: (2025)
GeoPos: A Minimal Positional Encoding for Enhanced Fine-Grained Details in Image Synthesis Using Convolutional Neural Networks
von: Hosseini, Mehran, et al.
Veröffentlicht: (2024)
von: Hosseini, Mehran, et al.
Veröffentlicht: (2024)
Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
von: Chandra, Abhranil, et al.
Veröffentlicht: (2025)
von: Chandra, Abhranil, et al.
Veröffentlicht: (2025)
SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues
von: Kuo, Martin, et al.
Veröffentlicht: (2025)
von: Kuo, Martin, et al.
Veröffentlicht: (2025)
Making New Connections: LLMs as Puzzle Generators for The New York Times' Connections Word Game
von: Merino, Tim, et al.
Veröffentlicht: (2024)
von: Merino, Tim, et al.
Veröffentlicht: (2024)
DialogueReason: Rule-Based RL Sparks Dialogue Reasoning in LLMs
von: Shu, Yubo, et al.
Veröffentlicht: (2025)
von: Shu, Yubo, et al.
Veröffentlicht: (2025)
Don't Forget to Connect! Improving RAG with Graph-based Reranking
von: Dong, Jialin, et al.
Veröffentlicht: (2024)
von: Dong, Jialin, et al.
Veröffentlicht: (2024)
Plantain: Plan-Answer Interleaved Reasoning
von: Liang, Anthony, et al.
Veröffentlicht: (2025)
von: Liang, Anthony, et al.
Veröffentlicht: (2025)
Well, that escalated quickly: The Single-Turn Crescendo Attack (STCA)
von: Aqrawi, Alan, et al.
Veröffentlicht: (2024)
von: Aqrawi, Alan, et al.
Veröffentlicht: (2024)
Not All LLM Reasoners Are Created Equal
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
von: Didolkar, Aniket, et al.
Veröffentlicht: (2025)
von: Didolkar, Aniket, et al.
Veröffentlicht: (2025)
IOLBENCH: Benchmarking LLMs on Linguistic Reasoning
von: Goyal, Satyam, et al.
Veröffentlicht: (2025)
von: Goyal, Satyam, et al.
Veröffentlicht: (2025)
Beyond Continuity: Challenges of Context Switching in Multi-Turn Dialogue with LLMs
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors
von: Wang, Jian, et al.
Veröffentlicht: (2025)
von: Wang, Jian, et al.
Veröffentlicht: (2025)
Off-Trajectory Reasoning: Can LLMs Collaborate on Reasoning Trajectory?
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
LLM-Driven Multi-Turn Task-Oriented Dialogue Synthesis for Realistic Reasoning
von: Zhu, Yu, et al.
Veröffentlicht: (2026)
von: Zhu, Yu, et al.
Veröffentlicht: (2026)
CARE: Turning LLMs Into Causal Reasoning Expert
von: Dong, Juncheng, et al.
Veröffentlicht: (2025)
von: Dong, Juncheng, et al.
Veröffentlicht: (2025)
Clinical Semantic Intelligence (CSI): Emulating the Cognitive Framework of the Expert Clinician for Comprehensive Oral Disease Diagnosis
von: Mashayekhi, Mohammad, et al.
Veröffentlicht: (2025)
von: Mashayekhi, Mohammad, et al.
Veröffentlicht: (2025)
Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades
von: Pona, Edoardo, et al.
Veröffentlicht: (2026)
von: Pona, Edoardo, et al.
Veröffentlicht: (2026)
Cost-Effective Attention Mechanisms for Low Resource Settings: Necessity & Sufficiency of Linear Transformations
von: Hosseini, Peyman, et al.
Veröffentlicht: (2024)
von: Hosseini, Peyman, et al.
Veröffentlicht: (2024)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
von: Jin, Keyan, et al.
Veröffentlicht: (2025)
von: Jin, Keyan, et al.
Veröffentlicht: (2025)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
Nonclassical Cutoff Fluctuations in Squeezed-Light-Driven High-Harmonic Generation
von: Khurelbaatar, Tsendsuren, et al.
Veröffentlicht: (2026)
von: Khurelbaatar, Tsendsuren, et al.
Veröffentlicht: (2026)
Strong-Field Quantum Metrology Beyond the Standard Quantum Limit
von: Khurelbaatar, Tsendsuren, et al.
Veröffentlicht: (2026)
von: Khurelbaatar, Tsendsuren, et al.
Veröffentlicht: (2026)
Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations
von: Liu, Zijie, et al.
Veröffentlicht: (2025)
von: Liu, Zijie, et al.
Veröffentlicht: (2025)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
von: Munkhdalai, Tsendsuren, et al.
Veröffentlicht: (2024) -
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2025) -
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
von: Fatemi, Bahare, et al.
Veröffentlicht: (2024) -
Let Your Graph Do the Talking: Encoding Structured Data for LLMs
von: Perozzi, Bryan, et al.
Veröffentlicht: (2024) -
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)