Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information
Fuente:
arXiv
Saved in:
| Main Authors: | Yim, Yauwai, Chan, Chunkit, Shi, Tianyu, Deng, Zheye, Fan, Wei, Zheng, Tianshi, Song, Yangqiu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
by: Liang, Fangzhou, et al.
Published: (2025)
by: Liang, Fangzhou, et al.
Published: (2025)
Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction
by: Deng, Zheye, et al.
Published: (2024)
by: Deng, Zheye, et al.
Published: (2024)
NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding
by: Chan, Chunkit, et al.
Published: (2024)
by: Chan, Chunkit, et al.
Published: (2024)
Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
by: Deng, Zheye, et al.
Published: (2025)
by: Deng, Zheye, et al.
Published: (2025)
DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
by: Mo, Yunxiang, et al.
Published: (2025)
by: Mo, Yunxiang, et al.
Published: (2025)
Constrained Reasoning Chains for Enhancing Theory-of-Mind in Large Language Models
by: Lin, Zizheng, et al.
Published: (2024)
by: Lin, Zizheng, et al.
Published: (2024)
Persona Knowledge-Aligned Prompt Tuning Method for Online Debate
by: Chan, Chunkit, et al.
Published: (2024)
by: Chan, Chunkit, et al.
Published: (2024)
XToM: Exploring the Multilingual Theory of Mind for Large Language Models
by: Chan, Chunkit, et al.
Published: (2025)
by: Chan, Chunkit, et al.
Published: (2025)
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
by: Zheng, Tianshi, et al.
Published: (2024)
by: Zheng, Tianshi, et al.
Published: (2024)
Enhancing Commentary Strategies for Imperfect Information Card Games: A Study of Large Language Models in Guandan Commentary
by: Tao, Meiling, et al.
Published: (2024)
by: Tao, Meiling, et al.
Published: (2024)
INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling
by: Shi, Haochen, et al.
Published: (2025)
by: Shi, Haochen, et al.
Published: (2025)
GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
by: Fan, Wei, et al.
Published: (2024)
by: Fan, Wei, et al.
Published: (2024)
Enhancing Transformers for Generalizable First-Order Logical Entailment
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities
by: Su, Ying, et al.
Published: (2024)
by: Su, Ying, et al.
Published: (2024)
Patterns Over Principles: The Fragility of Inductive Reasoning in LLMs under Noisy Observations
by: Li, Chunyang, et al.
Published: (2025)
by: Li, Chunyang, et al.
Published: (2025)
From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
Combining Theory of Mind and Abduction for Cooperation under Imperfect Information
by: Montes, Nieves, et al.
Published: (2022)
by: Montes, Nieves, et al.
Published: (2022)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
by: Zong, Qing, et al.
Published: (2024)
by: Zong, Qing, et al.
Published: (2024)
Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation
by: Wang, Jerry, et al.
Published: (2025)
by: Wang, Jerry, et al.
Published: (2025)
Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
Legal Rule Induction: Towards Generalizable Principle Discovery from Analogous Judicial Precedents
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
by: Zong, Qing, et al.
Published: (2025)
by: Zong, Qing, et al.
Published: (2025)
EventGround: Narrative Reasoning by Grounding to Eventuality-centric Knowledge Graphs
by: Jiayang, Cheng, et al.
Published: (2024)
by: Jiayang, Cheng, et al.
Published: (2024)
Mastering the Game of Guandan with Deep Reinforcement Learning and Behavior Regulating
by: Yanggong, Yifan, et al.
Published: (2024)
by: Yanggong, Yifan, et al.
Published: (2024)
Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4
by: Guo, Jiaxian, et al.
Published: (2023)
by: Guo, Jiaxian, et al.
Published: (2023)
InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding
by: Jiayang, Cheng, et al.
Published: (2025)
by: Jiayang, Cheng, et al.
Published: (2025)
ChatGPT Evaluation on Sentence Level Relations: A Focus on Temporal, Causal, and Discourse Relations
by: Chan, Chunkit, et al.
Published: (2023)
by: Chan, Chunkit, et al.
Published: (2023)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
Towards Multi-Agent Reasoning Systems for Collaborative Expertise Delegation: An Exploratory Design Study
by: Xu, Baixuan, et al.
Published: (2025)
by: Xu, Baixuan, et al.
Published: (2025)
SessionIntentBench: A Multi-task Inter-session Intention-shift Modeling Benchmark for E-commerce Customer Behavior Understanding
by: Yang, Yuqi, et al.
Published: (2025)
by: Yang, Yuqi, et al.
Published: (2025)
PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
by: Li, Haoran, et al.
Published: (2023)
by: Li, Haoran, et al.
Published: (2023)
KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education?
by: Zheng, Tianshi, et al.
Published: (2024)
by: Zheng, Tianshi, et al.
Published: (2024)
United Minds or Isolated Agents? Exploring Coordination of LLMs under Cognitive Load Theory
by: Shang, HaoYang, et al.
Published: (2025)
by: Shang, HaoYang, et al.
Published: (2025)
Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning
by: Fan, Wei, et al.
Published: (2026)
by: Fan, Wei, et al.
Published: (2026)
Cooperative Bayesian Optimization for Imperfect Agents
by: Khoshvishkaie, Ali, et al.
Published: (2024)
by: Khoshvishkaie, Ali, et al.
Published: (2024)
MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Digital Player: Evaluating Large Language Models based Human-like Agent in Games
by: Wang, Jiawei, et al.
Published: (2025)
by: Wang, Jiawei, et al.
Published: (2025)
Induced Representations in Cooperative Games with Homogeneous Groups of Players
by: Kiang, Windsor
Published: (2026)
by: Kiang, Windsor
Published: (2026)
Game Theory with Simulation of Other Players
by: Kovarik, Vojtech, et al.
Published: (2023)
by: Kovarik, Vojtech, et al.
Published: (2023)
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
by: Xu, Baixuan, et al.
Published: (2025)
by: Xu, Baixuan, et al.
Published: (2025)
Similar Items
-
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
by: Liang, Fangzhou, et al.
Published: (2025) -
Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction
by: Deng, Zheye, et al.
Published: (2024) -
NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding
by: Chan, Chunkit, et al.
Published: (2024) -
Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
by: Deng, Zheye, et al.
Published: (2025) -
DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
by: Mo, Yunxiang, et al.
Published: (2025)