Pencil Puzzle Bench: A Benchmark for Multi-Step Verifiable Reasoning
Fuente:
arXiv
Salvato in:
| Autore principale: | Waugh, Justin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
di: Li, Sijia, et al.
Pubblicazione: (2026)
di: Li, Sijia, et al.
Pubblicazione: (2026)
Monopoly Deal: A Benchmark Environment for Bounded One-Sided Response Games
di: Wolf, Will
Pubblicazione: (2025)
di: Wolf, Will
Pubblicazione: (2025)
Learning and Collusion in Multi-unit Auctions
di: Brânzei, Simina, et al.
Pubblicazione: (2023)
di: Brânzei, Simina, et al.
Pubblicazione: (2023)
PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations
di: Lei, Yingjie
Pubblicazione: (2026)
di: Lei, Yingjie
Pubblicazione: (2026)
NePPO: Near-Potential Policy Optimization for General-Sum Multi-Agent Reinforcement Learning
di: Kalanther, Addison, et al.
Pubblicazione: (2026)
di: Kalanther, Addison, et al.
Pubblicazione: (2026)
Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models
di: Hennes, Daniel, et al.
Pubblicazione: (2026)
di: Hennes, Daniel, et al.
Pubblicazione: (2026)
GemNet: Menu-Based, Strategy-Proof Multi-Bidder Auctions Through Deep Learning
di: Wang, Tonghan, et al.
Pubblicazione: (2024)
di: Wang, Tonghan, et al.
Pubblicazione: (2024)
LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
di: Jia, Jingru, et al.
Pubblicazione: (2025)
di: Jia, Jingru, et al.
Pubblicazione: (2025)
Multi-Head Attention Is a Multi-Player Game
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2026)
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2026)
Bandits with Preference Feedback: A Stackelberg Game Perspective
di: Pásztor, Barna, et al.
Pubblicazione: (2024)
di: Pásztor, Barna, et al.
Pubblicazione: (2024)
A Simple, Solid, and Reproducible Baseline for Bridge Bidding AI
di: Kita, Haruka, et al.
Pubblicazione: (2024)
di: Kita, Haruka, et al.
Pubblicazione: (2024)
A Framework for Adversarial Analysis of Decision Support Systems Prior to Deployment
di: Bissey, Brett, et al.
Pubblicazione: (2025)
di: Bissey, Brett, et al.
Pubblicazione: (2025)
SpinGPT: A Large-Language-Model Approach to Playing Poker Correctly
di: Maugin, Narada, et al.
Pubblicazione: (2025)
di: Maugin, Narada, et al.
Pubblicazione: (2025)
Optimizing Hard-to-Place Kidney Allocation: A Machine Learning Approach to Center Ranking
di: Berry, Sean, et al.
Pubblicazione: (2024)
di: Berry, Sean, et al.
Pubblicazione: (2024)
A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
di: Liu, Mingyang, et al.
Pubblicazione: (2024)
di: Liu, Mingyang, et al.
Pubblicazione: (2024)
Do LLM Agents Have Regret? A Case Study in Online Learning and Games
di: Park, Chanwoo, et al.
Pubblicazione: (2024)
di: Park, Chanwoo, et al.
Pubblicazione: (2024)
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers
di: Sun, Haoran, et al.
Pubblicazione: (2025)
di: Sun, Haoran, et al.
Pubblicazione: (2025)
ElementaryNet: A Non-Strategic Neural Network for Predicting Human Behavior in Normal-Form Games
di: d'Eon, Greg, et al.
Pubblicazione: (2025)
di: d'Eon, Greg, et al.
Pubblicazione: (2025)
Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game
di: Karabag, Mustafa O., et al.
Pubblicazione: (2025)
di: Karabag, Mustafa O., et al.
Pubblicazione: (2025)
Nash CoT: Multi-Path Inference with Preference Equilibrium
di: Zhang, Ziqi, et al.
Pubblicazione: (2024)
di: Zhang, Ziqi, et al.
Pubblicazione: (2024)
The Intelligent Disobedience Game: Formulating Disobedience in Stackelberg Games and Markov Decision Processes
di: Hornig, Benedikt, et al.
Pubblicazione: (2026)
di: Hornig, Benedikt, et al.
Pubblicazione: (2026)
Efficient Ensemble Selection from Binary and Pairwise Feedback
di: Neoh, Tzeh Yuan, et al.
Pubblicazione: (2026)
di: Neoh, Tzeh Yuan, et al.
Pubblicazione: (2026)
GroupSegment-SHAP: Shapley Value Explanations with Group-Segment Players for Multivariate Time Series
di: Kim, Jinwoong, et al.
Pubblicazione: (2026)
di: Kim, Jinwoong, et al.
Pubblicazione: (2026)
Finding Common Ground in a Sea of Alternatives
di: Chooi, Jay, et al.
Pubblicazione: (2026)
di: Chooi, Jay, et al.
Pubblicazione: (2026)
Asymmetric regularization mechanism for GAN training with Variational Inequalities
di: Giagtzoglou, Spyridon C., et al.
Pubblicazione: (2026)
di: Giagtzoglou, Spyridon C., et al.
Pubblicazione: (2026)
Strategic Candidacy in Generative AI Arenas
di: Hays, Chris, et al.
Pubblicazione: (2026)
di: Hays, Chris, et al.
Pubblicazione: (2026)
Adaptive Contracts for Cost-Effective AI Delegation
di: Saig, Eden, et al.
Pubblicazione: (2026)
di: Saig, Eden, et al.
Pubblicazione: (2026)
Governing AI Forgetting: Auditing for Machine Unlearning Compliance
di: Lin, Qinqi, et al.
Pubblicazione: (2026)
di: Lin, Qinqi, et al.
Pubblicazione: (2026)
Next-Token Prediction and Regret Minimization
di: Mohri, Mehryar, et al.
Pubblicazione: (2026)
di: Mohri, Mehryar, et al.
Pubblicazione: (2026)
The Optimal Sample Complexity of Linear Contracts
di: Høgsgaard, Mikael Møller
Pubblicazione: (2026)
di: Høgsgaard, Mikael Møller
Pubblicazione: (2026)
The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
di: La Malfa, Gabriele, et al.
Pubblicazione: (2026)
di: La Malfa, Gabriele, et al.
Pubblicazione: (2026)
Routing, Cascades, and User Choice for LLMs
di: Mahmood, Rafid
Pubblicazione: (2026)
di: Mahmood, Rafid
Pubblicazione: (2026)
Knowledge-Free Correlated Agreement for Incentivizing Federated Learning
di: Witt, Leon, et al.
Pubblicazione: (2026)
di: Witt, Leon, et al.
Pubblicazione: (2026)
In-Context Credit Assignment via the Core
di: Harris, Keegan, et al.
Pubblicazione: (2026)
di: Harris, Keegan, et al.
Pubblicazione: (2026)
When Individually Calibrated Models Become Collectively Miscalibrated
di: Wang, Zhaohui
Pubblicazione: (2026)
di: Wang, Zhaohui
Pubblicazione: (2026)
Ranking Abuse via Strategic Pairwise Data Perturbations
di: Yao, Junyi, et al.
Pubblicazione: (2026)
di: Yao, Junyi, et al.
Pubblicazione: (2026)
Real-Time Parallel Counterfactual Regret Minimization
di: Li, Boning, et al.
Pubblicazione: (2026)
di: Li, Boning, et al.
Pubblicazione: (2026)
Sharp Spectral Thresholds for Logit Fixed Points
di: Wang, Tongxi
Pubblicazione: (2026)
di: Wang, Tongxi
Pubblicazione: (2026)
Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Context Variables
di: Chen, Yang, et al.
Pubblicazione: (2025)
di: Chen, Yang, et al.
Pubblicazione: (2025)
LiteEFG: An Efficient Python Library for Solving Extensive-form Games
di: Liu, Mingyang, et al.
Pubblicazione: (2024)
di: Liu, Mingyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
di: Li, Sijia, et al.
Pubblicazione: (2026) -
Monopoly Deal: A Benchmark Environment for Bounded One-Sided Response Games
di: Wolf, Will
Pubblicazione: (2025) -
Learning and Collusion in Multi-unit Auctions
di: Brânzei, Simina, et al.
Pubblicazione: (2023) -
PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations
di: Lei, Yingjie
Pubblicazione: (2026) -
NePPO: Near-Potential Policy Optimization for General-Sum Multi-Agent Reinforcement Learning
di: Kalanther, Addison, et al.
Pubblicazione: (2026)