Pencil Puzzle Bench: A Benchmark for Multi-Step Verifiable Reasoning
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Waugh, Justin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
von: Li, Sijia, et al.
Veröffentlicht: (2026)
von: Li, Sijia, et al.
Veröffentlicht: (2026)
Monopoly Deal: A Benchmark Environment for Bounded One-Sided Response Games
von: Wolf, Will
Veröffentlicht: (2025)
von: Wolf, Will
Veröffentlicht: (2025)
Learning and Collusion in Multi-unit Auctions
von: Brânzei, Simina, et al.
Veröffentlicht: (2023)
von: Brânzei, Simina, et al.
Veröffentlicht: (2023)
PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations
von: Lei, Yingjie
Veröffentlicht: (2026)
von: Lei, Yingjie
Veröffentlicht: (2026)
NePPO: Near-Potential Policy Optimization for General-Sum Multi-Agent Reinforcement Learning
von: Kalanther, Addison, et al.
Veröffentlicht: (2026)
von: Kalanther, Addison, et al.
Veröffentlicht: (2026)
Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models
von: Hennes, Daniel, et al.
Veröffentlicht: (2026)
von: Hennes, Daniel, et al.
Veröffentlicht: (2026)
GemNet: Menu-Based, Strategy-Proof Multi-Bidder Auctions Through Deep Learning
von: Wang, Tonghan, et al.
Veröffentlicht: (2024)
von: Wang, Tonghan, et al.
Veröffentlicht: (2024)
LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
von: Jia, Jingru, et al.
Veröffentlicht: (2025)
von: Jia, Jingru, et al.
Veröffentlicht: (2025)
Multi-Head Attention Is a Multi-Player Game
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2026)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2026)
Bandits with Preference Feedback: A Stackelberg Game Perspective
von: Pásztor, Barna, et al.
Veröffentlicht: (2024)
von: Pásztor, Barna, et al.
Veröffentlicht: (2024)
A Simple, Solid, and Reproducible Baseline for Bridge Bidding AI
von: Kita, Haruka, et al.
Veröffentlicht: (2024)
von: Kita, Haruka, et al.
Veröffentlicht: (2024)
A Framework for Adversarial Analysis of Decision Support Systems Prior to Deployment
von: Bissey, Brett, et al.
Veröffentlicht: (2025)
von: Bissey, Brett, et al.
Veröffentlicht: (2025)
SpinGPT: A Large-Language-Model Approach to Playing Poker Correctly
von: Maugin, Narada, et al.
Veröffentlicht: (2025)
von: Maugin, Narada, et al.
Veröffentlicht: (2025)
Optimizing Hard-to-Place Kidney Allocation: A Machine Learning Approach to Center Ranking
von: Berry, Sean, et al.
Veröffentlicht: (2024)
von: Berry, Sean, et al.
Veröffentlicht: (2024)
A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
von: Liu, Mingyang, et al.
Veröffentlicht: (2024)
von: Liu, Mingyang, et al.
Veröffentlicht: (2024)
Do LLM Agents Have Regret? A Case Study in Online Learning and Games
von: Park, Chanwoo, et al.
Veröffentlicht: (2024)
von: Park, Chanwoo, et al.
Veröffentlicht: (2024)
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
ElementaryNet: A Non-Strategic Neural Network for Predicting Human Behavior in Normal-Form Games
von: d'Eon, Greg, et al.
Veröffentlicht: (2025)
von: d'Eon, Greg, et al.
Veröffentlicht: (2025)
Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game
von: Karabag, Mustafa O., et al.
Veröffentlicht: (2025)
von: Karabag, Mustafa O., et al.
Veröffentlicht: (2025)
Nash CoT: Multi-Path Inference with Preference Equilibrium
von: Zhang, Ziqi, et al.
Veröffentlicht: (2024)
von: Zhang, Ziqi, et al.
Veröffentlicht: (2024)
The Intelligent Disobedience Game: Formulating Disobedience in Stackelberg Games and Markov Decision Processes
von: Hornig, Benedikt, et al.
Veröffentlicht: (2026)
von: Hornig, Benedikt, et al.
Veröffentlicht: (2026)
Efficient Ensemble Selection from Binary and Pairwise Feedback
von: Neoh, Tzeh Yuan, et al.
Veröffentlicht: (2026)
von: Neoh, Tzeh Yuan, et al.
Veröffentlicht: (2026)
GroupSegment-SHAP: Shapley Value Explanations with Group-Segment Players for Multivariate Time Series
von: Kim, Jinwoong, et al.
Veröffentlicht: (2026)
von: Kim, Jinwoong, et al.
Veröffentlicht: (2026)
Finding Common Ground in a Sea of Alternatives
von: Chooi, Jay, et al.
Veröffentlicht: (2026)
von: Chooi, Jay, et al.
Veröffentlicht: (2026)
Asymmetric regularization mechanism for GAN training with Variational Inequalities
von: Giagtzoglou, Spyridon C., et al.
Veröffentlicht: (2026)
von: Giagtzoglou, Spyridon C., et al.
Veröffentlicht: (2026)
Strategic Candidacy in Generative AI Arenas
von: Hays, Chris, et al.
Veröffentlicht: (2026)
von: Hays, Chris, et al.
Veröffentlicht: (2026)
Adaptive Contracts for Cost-Effective AI Delegation
von: Saig, Eden, et al.
Veröffentlicht: (2026)
von: Saig, Eden, et al.
Veröffentlicht: (2026)
Governing AI Forgetting: Auditing for Machine Unlearning Compliance
von: Lin, Qinqi, et al.
Veröffentlicht: (2026)
von: Lin, Qinqi, et al.
Veröffentlicht: (2026)
Next-Token Prediction and Regret Minimization
von: Mohri, Mehryar, et al.
Veröffentlicht: (2026)
von: Mohri, Mehryar, et al.
Veröffentlicht: (2026)
The Optimal Sample Complexity of Linear Contracts
von: Høgsgaard, Mikael Møller
Veröffentlicht: (2026)
von: Høgsgaard, Mikael Møller
Veröffentlicht: (2026)
The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
von: La Malfa, Gabriele, et al.
Veröffentlicht: (2026)
von: La Malfa, Gabriele, et al.
Veröffentlicht: (2026)
Routing, Cascades, and User Choice for LLMs
von: Mahmood, Rafid
Veröffentlicht: (2026)
von: Mahmood, Rafid
Veröffentlicht: (2026)
Knowledge-Free Correlated Agreement for Incentivizing Federated Learning
von: Witt, Leon, et al.
Veröffentlicht: (2026)
von: Witt, Leon, et al.
Veröffentlicht: (2026)
In-Context Credit Assignment via the Core
von: Harris, Keegan, et al.
Veröffentlicht: (2026)
von: Harris, Keegan, et al.
Veröffentlicht: (2026)
When Individually Calibrated Models Become Collectively Miscalibrated
von: Wang, Zhaohui
Veröffentlicht: (2026)
von: Wang, Zhaohui
Veröffentlicht: (2026)
Ranking Abuse via Strategic Pairwise Data Perturbations
von: Yao, Junyi, et al.
Veröffentlicht: (2026)
von: Yao, Junyi, et al.
Veröffentlicht: (2026)
Real-Time Parallel Counterfactual Regret Minimization
von: Li, Boning, et al.
Veröffentlicht: (2026)
von: Li, Boning, et al.
Veröffentlicht: (2026)
Sharp Spectral Thresholds for Logit Fixed Points
von: Wang, Tongxi
Veröffentlicht: (2026)
von: Wang, Tongxi
Veröffentlicht: (2026)
Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Context Variables
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
LiteEFG: An Efficient Python Library for Solving Extensive-form Games
von: Liu, Mingyang, et al.
Veröffentlicht: (2024)
von: Liu, Mingyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
von: Li, Sijia, et al.
Veröffentlicht: (2026) -
Monopoly Deal: A Benchmark Environment for Bounded One-Sided Response Games
von: Wolf, Will
Veröffentlicht: (2025) -
Learning and Collusion in Multi-unit Auctions
von: Brânzei, Simina, et al.
Veröffentlicht: (2023) -
PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations
von: Lei, Yingjie
Veröffentlicht: (2026) -
NePPO: Near-Potential Policy Optimization for General-Sum Multi-Agent Reinforcement Learning
von: Kalanther, Addison, et al.
Veröffentlicht: (2026)