DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
Fuente:
arXiv
Saved in:
| Main Authors: | Mo, Yunxiang, Zheng, Tianshi, Zong, Qing, Liu, Jiayu, Xu, Baixuan, Yim, Yauwai, Chan, Chunkit, Bai, Jiaxin, Song, Yangqiu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
by: Liang, Fangzhou, et al.
Published: (2025)
by: Liang, Fangzhou, et al.
Published: (2025)
Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information
by: Yim, Yauwai, et al.
Published: (2024)
by: Yim, Yauwai, et al.
Published: (2024)
Dixit
Published: (2019)
Published: (2019)
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
by: Zheng, Tianshi, et al.
Published: (2024)
by: Zheng, Tianshi, et al.
Published: (2024)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction
by: Deng, Zheye, et al.
Published: (2024)
by: Deng, Zheye, et al.
Published: (2024)
Advancing Abductive Reasoning in Knowledge Graphs through Complex Logical Hypothesis Generation
by: Bai, Jiaxin, et al.
Published: (2023)
by: Bai, Jiaxin, et al.
Published: (2023)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
by: Zong, Qing, et al.
Published: (2025)
by: Zong, Qing, et al.
Published: (2025)
Evaluating recipes for development success / Avinash Dixit
by: Dixit, Avinash
Published: (2006)
by: Dixit, Avinash
Published: (2006)
Dixit. Gra filozoficzna zamiast gamingu
by: Mariusz Mazurkiewicz
Published: (2013)
by: Mariusz Mazurkiewicz
Published: (2013)
Controllable Logical Hypothesis Generation for Abductive Reasoning in Knowledge Graphs
by: Gao, Yisen, et al.
Published: (2025)
by: Gao, Yisen, et al.
Published: (2025)
Optimization in economic theory / Avinash K. Dixit
by: Dixit, Avinash K
by: Dixit, Avinash K
Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
by: Deng, Zheye, et al.
Published: (2025)
by: Deng, Zheye, et al.
Published: (2025)
LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
Unifying Deductive and Abductive Reasoning in Knowledge Graphs with Masked Diffusion Model
by: Gao, Yisen, et al.
Published: (2025)
by: Gao, Yisen, et al.
Published: (2025)
On the Double Lambert Series Conjecture of Andrews-Dixit--Schultz-Yee
by: Fang, Qianwen
Published: (2026)
by: Fang, Qianwen
Published: (2026)
Constrained Reasoning Chains for Enhancing Theory-of-Mind in Large Language Models
by: Lin, Zizheng, et al.
Published: (2024)
by: Lin, Zizheng, et al.
Published: (2024)
INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling
by: Shi, Haochen, et al.
Published: (2025)
by: Shi, Haochen, et al.
Published: (2025)
NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding
by: Chan, Chunkit, et al.
Published: (2024)
by: Chan, Chunkit, et al.
Published: (2024)
Persona Knowledge-Aligned Prompt Tuning Method for Online Debate
by: Chan, Chunkit, et al.
Published: (2024)
by: Chan, Chunkit, et al.
Published: (2024)
CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense Reasoning
by: Wang, Weiqi, et al.
Published: (2024)
by: Wang, Weiqi, et al.
Published: (2024)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
by: Zong, Qing, et al.
Published: (2024)
by: Zong, Qing, et al.
Published: (2024)
ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities
by: Su, Ying, et al.
Published: (2024)
by: Su, Ying, et al.
Published: (2024)
AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
by: Tsang, Hong Ting, et al.
Published: (2025)
by: Tsang, Hong Ting, et al.
Published: (2025)
Towards Multi-Agent Reasoning Systems for Collaborative Expertise Delegation: An Exploratory Design Study
by: Xu, Baixuan, et al.
Published: (2025)
by: Xu, Baixuan, et al.
Published: (2025)
EventGround: Narrative Reasoning by Grounding to Eventuality-centric Knowledge Graphs
by: Jiayang, Cheng, et al.
Published: (2024)
by: Jiayang, Cheng, et al.
Published: (2024)
Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
India y Pakist n. Las medidas necesarias de creación de confianza / Aabha Dixit
by: Dixit, Aabha
Published: (1996)
by: Dixit, Aabha
Published: (1996)
Interfaces naturales para contenidos digitales interactivos en museos. La experiencia de Galicia Dixital.
by: Luis Hernández Ibañez
Published: (2008)
by: Luis Hernández Ibañez
Published: (2008)
KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education?
by: Zheng, Tianshi, et al.
Published: (2024)
by: Zheng, Tianshi, et al.
Published: (2024)
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
by: Xu, Baixuan, et al.
Published: (2025)
by: Xu, Baixuan, et al.
Published: (2025)
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
by: Liu, Jiayu, et al.
Published: (2026)
by: Liu, Jiayu, et al.
Published: (2026)
The art of strategy : a game theorist's guide to success in business and life / Avinash K. Dixit, Barry J. Nalebuff
by: Dixit, Avinash K
Published: (2010)
by: Dixit, Avinash K
Published: (2010)
Thinking Strategically : the Competitive Edge un Business, Politic, and Everyday Life / vinask K. Dixit, Barry J. Nalebuff
by: Dixit, Avinash
Published: (1991)
by: Dixit, Avinash
Published: (1991)
The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
EcomEdit: An Automated E-commerce Knowledge Editing Framework for Enhanced Product and Purchase Intention Understanding
by: Lau, Ching Ming Samuel, et al.
Published: (2024)
by: Lau, Ching Ming Samuel, et al.
Published: (2024)
InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding
by: Jiayang, Cheng, et al.
Published: (2025)
by: Jiayang, Cheng, et al.
Published: (2025)
Patterns Over Principles: The Fragility of Inductive Reasoning in LLMs under Noisy Observations
by: Li, Chunyang, et al.
Published: (2025)
by: Li, Chunyang, et al.
Published: (2025)
ChatGPT Evaluation on Sentence Level Relations: A Focus on Temporal, Causal, and Discourse Relations
by: Chan, Chunkit, et al.
Published: (2023)
by: Chan, Chunkit, et al.
Published: (2023)
SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision
by: Liu, Yuxuan, et al.
Published: (2026)
by: Liu, Yuxuan, et al.
Published: (2026)
Similar Items
-
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
by: Liang, Fangzhou, et al.
Published: (2025) -
Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information
by: Yim, Yauwai, et al.
Published: (2024) -
Dixit
Published: (2019) -
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
by: Zheng, Tianshi, et al.
Published: (2024) -
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
by: Balepur, Nishant, et al.
Published: (2025)