SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Jianzhu, Wang, Kevin, Hsieh, Ryan, Zhou, Haisu, Zou, Tianqing, Cheng, Zerui, Wang, Zhangyang, Viswanath, Pramod |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VeRA: Verified Reasoning Data Augmentation at Scale
by: Cheng, Zerui, et al.
Published: (2026)
by: Cheng, Zerui, et al.
Published: (2026)
TabularMath: Evaluating Computational Extrapolation in Tabular Learning via Program-Verified Synthesis
by: Cheng, Zerui, et al.
Published: (2026)
by: Cheng, Zerui, et al.
Published: (2026)
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
by: Zheng, Zihan, et al.
Published: (2025)
by: Zheng, Zihan, et al.
Published: (2025)
TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks
by: Yao, Jianzhu, et al.
Published: (2025)
by: Yao, Jianzhu, et al.
Published: (2025)
MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games
by: Xie, Yunfei, et al.
Published: (2026)
by: Xie, Yunfei, et al.
Published: (2026)
Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
by: Kim, Kyuyoung, et al.
Published: (2026)
by: Kim, Kyuyoung, et al.
Published: (2026)
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
by: Wang, Kevin, et al.
Published: (2026)
by: Wang, Kevin, et al.
Published: (2026)
Split the Yield, Share the Risk: Pricing, Hedging and Fixed rates in DeFi
by: Nadkarni, Viraj, et al.
Published: (2025)
by: Nadkarni, Viraj, et al.
Published: (2025)
How Well Do LLMs Identify Cultural Unity in Diversity?
by: Li, Jialin, et al.
Published: (2024)
by: Li, Jialin, et al.
Published: (2024)
RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?
by: Dai, Yuyang, et al.
Published: (2026)
by: Dai, Yuyang, et al.
Published: (2026)
AutoCode: LLMs as Problem Setters for Competitive Programming
by: Zhou, Shang, et al.
Published: (2025)
by: Zhou, Shang, et al.
Published: (2025)
FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?
by: Wu, Eric, et al.
Published: (2024)
by: Wu, Eric, et al.
Published: (2024)
FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations
by: Hu, Xuesi, et al.
Published: (2026)
by: Hu, Xuesi, et al.
Published: (2026)
How Well Do LLMs Understand Tunisian Arabic?
by: Mahdi, Mohamed
Published: (2025)
by: Mahdi, Mohamed
Published: (2025)
How Well Do LLMs Understand Drug Mechanisms? A Knowledge + Reasoning Evaluation Dataset
by: Mohan, Sunil, et al.
Published: (2025)
by: Mohan, Sunil, et al.
Published: (2025)
How Well Do LLMs Perform on the Simplest Long-Chain Reasoning Tasks: An Empirical Study on the Equivalence Class Problem
by: Zheng, Chun, et al.
Published: (2026)
by: Zheng, Chun, et al.
Published: (2026)
CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models
by: Irwin, Lucas, et al.
Published: (2025)
by: Irwin, Lucas, et al.
Published: (2025)
AgileRate: Bringing Adaptivity and Robustness to DeFi Lending Markets
by: Bastankhah, Mahsa, et al.
Published: (2024)
by: Bastankhah, Mahsa, et al.
Published: (2024)
Adaptive Curves for Optimally Efficient Market Making
by: Nadkarni, Viraj, et al.
Published: (2024)
by: Nadkarni, Viraj, et al.
Published: (2024)
Training AI to be Loyal
by: Oh, Sewoong, et al.
Published: (2025)
by: Oh, Sewoong, et al.
Published: (2025)
How Well Do LLMs Imitate Human Writing Style?
by: Jemama, Rebira, et al.
Published: (2025)
by: Jemama, Rebira, et al.
Published: (2025)
How Likely Do LLMs with CoT Mimic Human Reasoning?
by: Bao, Guangsheng, et al.
Published: (2024)
by: Bao, Guangsheng, et al.
Published: (2024)
How Well Do Multimodal Models Reason on ECG Signals?
by: Xu, Maxwell A., et al.
Published: (2026)
by: Xu, Maxwell A., et al.
Published: (2026)
How Well Do Multi-modal LLMs Interpret CT Scans? An Auto-Evaluation Framework for Analyses
by: Zhu, Qingqing, et al.
Published: (2024)
by: Zhu, Qingqing, et al.
Published: (2024)
Understanding the Impact of Negative Prompts: When and How Do They Take Effect?
by: Ban, Yuanhao, et al.
Published: (2024)
by: Ban, Yuanhao, et al.
Published: (2024)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning
by: Wang, Peihao, et al.
Published: (2025)
by: Wang, Peihao, et al.
Published: (2025)
Do Persona-Infused LLMs Affect Performance in a Strategic Reasoning Game?
by: Licato, John, et al.
Published: (2025)
by: Licato, John, et al.
Published: (2025)
How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models
by: Jiang, Jiyue, et al.
Published: (2024)
by: Jiang, Jiyue, et al.
Published: (2024)
Do Compressed LLMs Forget Knowledge? An Experimental Study with Practical Implications
by: Hoang, Duc N. M, et al.
Published: (2023)
by: Hoang, Duc N. M, et al.
Published: (2023)
Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?
by: Pan, Zhenyu, et al.
Published: (2024)
by: Pan, Zhenyu, et al.
Published: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026)
by: Wang, Erchi, et al.
Published: (2026)
CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V
by: Chen, John, et al.
Published: (2026)
by: Chen, John, et al.
Published: (2026)
ParlayMarket: Automated Market Making for Parlay-style Joint Contracts
by: Rana, Ranvir, et al.
Published: (2026)
by: Rana, Ranvir, et al.
Published: (2026)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
by: Liu, Qing'an, et al.
Published: (2026)
by: Liu, Qing'an, et al.
Published: (2026)
SPIN AND RELATIVITY: A SEMICLASSICAL MODEL FOR ELECTRON SPIN
by: Héctor Torres-Silva
Published: (2008)
by: Héctor Torres-Silva
Published: (2008)
Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
by: He, Yidong, et al.
Published: (2026)
by: He, Yidong, et al.
Published: (2026)
Strategic Planning and Rationalizing on Trees Make LLMs Better Debaters
by: Wang, Danqing, et al.
Published: (2025)
by: Wang, Danqing, et al.
Published: (2025)
How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained Knowledge
by: Gao, Wayne, et al.
Published: (2026)
by: Gao, Wayne, et al.
Published: (2026)
Similar Items
-
VeRA: Verified Reasoning Data Augmentation at Scale
by: Cheng, Zerui, et al.
Published: (2026) -
TabularMath: Evaluating Computational Extrapolation in Tabular Learning via Program-Verified Synthesis
by: Cheng, Zerui, et al.
Published: (2026) -
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
by: Zheng, Zihan, et al.
Published: (2025) -
TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks
by: Yao, Jianzhu, et al.
Published: (2025) -
MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games
by: Xie, Yunfei, et al.
Published: (2026)