Towards Understanding Specification Gaming in Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Nishimura-Gasparian, Kei, McCarthy, Robert, Lindner, David |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Early Signs of Steganographic Capabilities in Frontier LLMs
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
Reasoning Models Struggle to Control their Chains of Thought
by: Yueh-Han, Chen, et al.
Published: (2026)
by: Yueh-Han, Chen, et al.
Published: (2026)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
by: Fronsdal, Kai, et al.
Published: (2024)
by: Fronsdal, Kai, et al.
Published: (2024)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
Building Understandable Messaging for Policy and Evidence Review (BUMPER) with AI
by: Rosenfeld, Katherine A., et al.
Published: (2024)
by: Rosenfeld, Katherine A., et al.
Published: (2024)
Evaluating and Understanding Scheming Propensity in LLM Agents
by: Hopman, Mia, et al.
Published: (2026)
by: Hopman, Mia, et al.
Published: (2026)
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
by: McCarthy, James, et al.
Published: (2025)
by: McCarthy, James, et al.
Published: (2025)
Compositional Understanding in Signaling Games
by: Freeborn, David Peter Wallis
Published: (2025)
by: Freeborn, David Peter Wallis
Published: (2025)
Towards Understanding the Cognitive Habits of Large Reasoning Models
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
A CMOS Probabilistic Computing Chip With In-situ hardware Aware Learning
by: Jhonsa, Jinesh, et al.
Published: (2025)
by: Jhonsa, Jinesh, et al.
Published: (2025)
Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical Tasks
by: Maeda, Kiyosu, et al.
Published: (2026)
by: Maeda, Kiyosu, et al.
Published: (2026)
Toward Modeling Player-Specific Chess Behaviors
by: Sogliuzzo, Loris, et al.
Published: (2026)
by: Sogliuzzo, Loris, et al.
Published: (2026)
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
by: Liao, Yi, et al.
Published: (2025)
by: Liao, Yi, et al.
Published: (2025)
Recontextualization Mitigates Specification Gaming without Modifying the Specification
by: Azarbal, Ariana, et al.
Published: (2025)
by: Azarbal, Ariana, et al.
Published: (2025)
Toward Formalizing LLM-Based Agent Designs through Structural Context Modeling and Semantic Dynamics Analysis
by: Jia, Haoyu, et al.
Published: (2026)
by: Jia, Haoyu, et al.
Published: (2026)
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
by: Piedrahita, David Guzman, et al.
Published: (2025)
by: Piedrahita, David Guzman, et al.
Published: (2025)
Towards a Mechanistic Understanding of Propositional Logical Reasoning in Large Language Models
by: Chen, Danchun, et al.
Published: (2026)
by: Chen, Danchun, et al.
Published: (2026)
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Toward the Axiomatization of Intelligence: Structure, Time, and Existence
by: Itoh, Kei
Published: (2025)
by: Itoh, Kei
Published: (2025)
The Token Games: Evaluating Language Model Reasoning with Puzzle Duels
by: Henniger, Simon, et al.
Published: (2026)
by: Henniger, Simon, et al.
Published: (2026)
Large language models can learn and generalize steganographic chain-of-thought under process supervision
by: Skaf, Joey, et al.
Published: (2025)
by: Skaf, Joey, et al.
Published: (2025)
Circular Reasoning: Understanding Self-Reinforcing Loops in Large Reasoning Models
by: Duan, Zenghao, et al.
Published: (2026)
by: Duan, Zenghao, et al.
Published: (2026)
Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability
by: Huang, Fan, et al.
Published: (2026)
by: Huang, Fan, et al.
Published: (2026)
Applying Natural Language Processing and Hierarchical Machine Learning Approaches to Text Difficulty Classification
by: Balyan, Renu, et al.
Published: (2020)
by: Balyan, Renu, et al.
Published: (2020)
Applying Natural Language Processing and Hierarchical Machine Learning Approaches to Text Difficulty Classification
by: Balyan, Renu, et al.
Published: (2020)
by: Balyan, Renu, et al.
Published: (2020)
Game On: Towards Language Models as RL Experimenters
by: Zhang, Jingwei, et al.
Published: (2024)
by: Zhang, Jingwei, et al.
Published: (2024)
Reasoning, Memorization, and Fine-Tuning Language Models for Non-Cooperative Games
by: Yang, Yunhao, et al.
Published: (2024)
by: Yang, Yunhao, et al.
Published: (2024)
CreativeGame:Toward Mechanic-Aware Creative Game Generation
by: Ma, Hongnan, et al.
Published: (2026)
by: Ma, Hongnan, et al.
Published: (2026)
Solving a Stackelberg Game on Transportation Networks in a Dynamic Crime Scenario: A Mixed Approach on Multi-Layer Networks
by: Samanta, Sukanya, et al.
Published: (2024)
by: Samanta, Sukanya, et al.
Published: (2024)
Enhance Reasoning for Large Language Models in the Game Werewolf
by: Wu, Shuang, et al.
Published: (2024)
by: Wu, Shuang, et al.
Published: (2024)
Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
by: He, Yidong, et al.
Published: (2026)
by: He, Yidong, et al.
Published: (2026)
Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
by: Wang, Pinzheng, et al.
Published: (2025)
by: Wang, Pinzheng, et al.
Published: (2025)
Limits of Emergent Reasoning of Large Language Models in Agentic Frameworks for Deterministic Games
by: Su, Chris, et al.
Published: (2025)
by: Su, Chris, et al.
Published: (2025)
Ludax: A GPU-Accelerated Domain Specific Language for Board Games
by: Todd, Graham, et al.
Published: (2025)
by: Todd, Graham, et al.
Published: (2025)
Internalizing Safety Understanding in Large Reasoning Models via Verification
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
The Geometry of Self-Verification in a Task-Specific Reasoning Model
by: Lee, Andrew, et al.
Published: (2025)
by: Lee, Andrew, et al.
Published: (2025)
Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games
by: Fan, Mingyuan, et al.
Published: (2026)
by: Fan, Mingyuan, et al.
Published: (2026)
HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
by: Liang, Jingcong, et al.
Published: (2025)
by: Liang, Jingcong, et al.
Published: (2025)
Respiratory Inhaler Sound Event Classification Using Self-Supervised Learning
by: Panah, Davoud Shariat, et al.
Published: (2025)
by: Panah, Davoud Shariat, et al.
Published: (2025)
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
by: Xu, Chenhui, et al.
Published: (2025)
by: Xu, Chenhui, et al.
Published: (2025)
Similar Items
-
Early Signs of Steganographic Capabilities in Frontier LLMs
by: Zolkowski, Artur, et al.
Published: (2025) -
Reasoning Models Struggle to Control their Chains of Thought
by: Yueh-Han, Chen, et al.
Published: (2026) -
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
by: Fronsdal, Kai, et al.
Published: (2024) -
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025) -
Building Understandable Messaging for Policy and Evidence Review (BUMPER) with AI
by: Rosenfeld, Katherine A., et al.
Published: (2024)