TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games
Fuente:
arXiv
Salvato in:
| Autori principali: | Yuan, Yuan, He, Muyu, Shahid, Muhammad Adil, Huang, Jiani, Li, Ziyang, Zhang, Li |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Vision Language Models Cannot Plan, but Can They Formalize?
di: He, Muyu, et al.
Pubblicazione: (2025)
di: He, Muyu, et al.
Pubblicazione: (2025)
Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
di: Song, Zirui, et al.
Pubblicazione: (2025)
di: Song, Zirui, et al.
Pubblicazione: (2025)
The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
di: He, Muyu, et al.
Pubblicazione: (2025)
di: He, Muyu, et al.
Pubblicazione: (2025)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
di: Jing, Huihao, et al.
Pubblicazione: (2025)
di: Jing, Huihao, et al.
Pubblicazione: (2025)
Toward Honest Language Models for Deductive Reasoning
di: Liu, Jiarui, et al.
Pubblicazione: (2025)
di: Liu, Jiarui, et al.
Pubblicazione: (2025)
LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
di: Zheng, Tianshi, et al.
Pubblicazione: (2025)
di: Zheng, Tianshi, et al.
Pubblicazione: (2025)
Hypothesis Testing Prompting Improves Deductive Reasoning in Large Language Models
di: Li, Yitian, et al.
Pubblicazione: (2024)
di: Li, Yitian, et al.
Pubblicazione: (2024)
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
di: Shi, Jiajun, et al.
Pubblicazione: (2025)
di: Shi, Jiajun, et al.
Pubblicazione: (2025)
How Far Are We from Intelligent Visual Deductive Reasoning?
di: Zhang, Yizhe, et al.
Pubblicazione: (2024)
di: Zhang, Yizhe, et al.
Pubblicazione: (2024)
Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns
di: Li, Xiang, et al.
Pubblicazione: (2025)
di: Li, Xiang, et al.
Pubblicazione: (2025)
Self-playing Adversarial Language Game Enhances LLM Reasoning
di: Cheng, Pengyu, et al.
Pubblicazione: (2024)
di: Cheng, Pengyu, et al.
Pubblicazione: (2024)
Could Thinking Multilingually Empower LLM Reasoning?
di: Gao, Changjiang, et al.
Pubblicazione: (2025)
di: Gao, Changjiang, et al.
Pubblicazione: (2025)
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
di: Chen, Michael K., et al.
Pubblicazione: (2025)
di: Chen, Michael K., et al.
Pubblicazione: (2025)
Understanding LLM Reasoning for Abstractive Summarization
di: Yuan, Haohan, et al.
Pubblicazione: (2025)
di: Yuan, Haohan, et al.
Pubblicazione: (2025)
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments
di: Tang, Wenjie, et al.
Pubblicazione: (2025)
di: Tang, Wenjie, et al.
Pubblicazione: (2025)
\texttt{ReMind}: Understanding Deductive Code Reasoning in LLMs
di: Gao, Jun, et al.
Pubblicazione: (2025)
di: Gao, Jun, et al.
Pubblicazione: (2025)
How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
The Role of Deductive and Inductive Reasoning in Large Language Models
di: Cai, Chengkun, et al.
Pubblicazione: (2024)
di: Cai, Chengkun, et al.
Pubblicazione: (2024)
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
di: Pandey, Atharva, et al.
Pubblicazione: (2025)
di: Pandey, Atharva, et al.
Pubblicazione: (2025)
Deductive Beam Search: Decoding Deducible Rationale for Chain-of-Thought Reasoning
di: Zhu, Tinghui, et al.
Pubblicazione: (2024)
di: Zhu, Tinghui, et al.
Pubblicazione: (2024)
GameArena: Evaluating LLM Reasoning through Live Computer Games
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
Revealing Algorithmic Deductive Circuits for Logical Reasoning
di: Nguyen, Phuong Minh, et al.
Pubblicazione: (2026)
di: Nguyen, Phuong Minh, et al.
Pubblicazione: (2026)
Leveraging LLMs for Hypothetical Deduction in Logical Inference: A Neuro-Symbolic Approach
di: Li, Qingchuan, et al.
Pubblicazione: (2024)
di: Li, Qingchuan, et al.
Pubblicazione: (2024)
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
di: Kim, Byungjun, et al.
Pubblicazione: (2024)
di: Kim, Byungjun, et al.
Pubblicazione: (2024)
MIND-Skill: Quality-Guaranteed Skill Generation via Multi-Agent Induction and Deduction
di: Li, Yixuan, et al.
Pubblicazione: (2026)
di: Li, Yixuan, et al.
Pubblicazione: (2026)
MindMerger: Efficient Boosting LLM Reasoning in non-English Languages
di: Huang, Zixian, et al.
Pubblicazione: (2024)
di: Huang, Zixian, et al.
Pubblicazione: (2024)
BioProBench: Comprehensive Dataset and Benchmark in Biological Protocol Understanding and Reasoning
di: Liu, Yuyang, et al.
Pubblicazione: (2025)
di: Liu, Yuyang, et al.
Pubblicazione: (2025)
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025)
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025)
Toward Mechanistic Explanation of Deductive Reasoning in Language Models
di: Maltoni, Davide, et al.
Pubblicazione: (2025)
di: Maltoni, Davide, et al.
Pubblicazione: (2025)
Investigating the Robustness of Deductive Reasoning with Large Language Models
di: Hoppe, Fabian, et al.
Pubblicazione: (2025)
di: Hoppe, Fabian, et al.
Pubblicazione: (2025)
EVEDIT: Event-based Knowledge Editing with Deductive Editing Boundaries
di: Liu, Jiateng, et al.
Pubblicazione: (2024)
di: Liu, Jiateng, et al.
Pubblicazione: (2024)
AMONGAGENTS: Evaluating Large Language Models in the Interactive Text-Based Social Deduction Game
di: Chi, Yizhou, et al.
Pubblicazione: (2024)
di: Chi, Yizhou, et al.
Pubblicazione: (2024)
Def-DTS: Deductive Reasoning for Open-domain Dialogue Topic Segmentation
di: Lee, Seungmin, et al.
Pubblicazione: (2025)
di: Lee, Seungmin, et al.
Pubblicazione: (2025)
Data Analysis and Performance Evaluation of Simulation Deduction Based on LLMs
di: Zhang, Shansi, et al.
Pubblicazione: (2025)
di: Zhang, Shansi, et al.
Pubblicazione: (2025)
Reinforced Internal-External Knowledge Synergistic Reasoning for Efficient Adaptive Search Agent
di: Huang, Ziyang, et al.
Pubblicazione: (2025)
di: Huang, Ziyang, et al.
Pubblicazione: (2025)
ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
di: He, Liyang, et al.
Pubblicazione: (2025)
di: He, Liyang, et al.
Pubblicazione: (2025)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
di: Sun, Jiaxing, et al.
Pubblicazione: (2024)
di: Sun, Jiaxing, et al.
Pubblicazione: (2024)
Software Engineering Methods For AI-Driven Deductive Legal Reasoning
di: Padhye, Rohan
Pubblicazione: (2024)
di: Padhye, Rohan
Pubblicazione: (2024)
GAMEBoT: Transparent Assessment of LLM Reasoning in Games
di: Lin, Wenye, et al.
Pubblicazione: (2024)
di: Lin, Wenye, et al.
Pubblicazione: (2024)
Rethinking Easy-to-Hard: Limits of Curriculum Learning in Post-Training for Deductive Reasoning
di: Mordig, Maximilian, et al.
Pubblicazione: (2026)
di: Mordig, Maximilian, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Vision Language Models Cannot Plan, but Can They Formalize?
di: He, Muyu, et al.
Pubblicazione: (2025) -
Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
di: Song, Zirui, et al.
Pubblicazione: (2025) -
The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
di: He, Muyu, et al.
Pubblicazione: (2025) -
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
di: Jing, Huihao, et al.
Pubblicazione: (2025) -
Toward Honest Language Models for Deductive Reasoning
di: Liu, Jiarui, et al.
Pubblicazione: (2025)