Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hwang, Dongyoon, Lee, Hojoon, Choo, Jaegul, Park, Dongmin, Park, Jongho |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
von: Kim, Hyunseung, et al.
Veröffentlicht: (2024)
von: Kim, Hyunseung, et al.
Veröffentlicht: (2024)
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
von: Kim, Donghu, et al.
Veröffentlicht: (2024)
von: Kim, Donghu, et al.
Veröffentlicht: (2024)
When Model Meets New Normals: Test-time Adaptation for Unsupervised Time-series Anomaly Detection
von: Kim, Dongmin, et al.
Veröffentlicht: (2023)
von: Kim, Dongmin, et al.
Veröffentlicht: (2023)
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning
von: Lee, Hojoon, et al.
Veröffentlicht: (2024)
von: Lee, Hojoon, et al.
Veröffentlicht: (2024)
ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2024)
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2024)
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
Self-Supervised Contrastive Learning for Long-term Forecasting
von: Park, Junwoo, et al.
Veröffentlicht: (2024)
von: Park, Junwoo, et al.
Veröffentlicht: (2024)
EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language Models
von: Kim, Jinhee, et al.
Veröffentlicht: (2024)
von: Kim, Jinhee, et al.
Veröffentlicht: (2024)
Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise Networks
von: Lee, Hojoon, et al.
Veröffentlicht: (2024)
von: Lee, Hojoon, et al.
Veröffentlicht: (2024)
Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target
von: Kim, Taesan, et al.
Veröffentlicht: (2026)
von: Kim, Taesan, et al.
Veröffentlicht: (2026)
Active Learning for Continual Learning: Keeping the Past Alive in the Present
von: Park, Jaehyun, et al.
Veröffentlicht: (2025)
von: Park, Jaehyun, et al.
Veröffentlicht: (2025)
ChessQA: Evaluating Large Language Models for Chess Understanding
von: Wen, Qianfeng, et al.
Veröffentlicht: (2025)
von: Wen, Qianfeng, et al.
Veröffentlicht: (2025)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
Domain-Adaptive Health Indicator Learning with Degradation-Stage Synchronized Sampling and Cross-Domain Autoencoder
von: Choo, Jungho, et al.
Veröffentlicht: (2026)
von: Choo, Jungho, et al.
Veröffentlicht: (2026)
Revisiting LLMs as Zero-Shot Time-Series Forecasters: Small Noise Can Break Large Models
von: Park, Junwoo, et al.
Veröffentlicht: (2025)
von: Park, Junwoo, et al.
Veröffentlicht: (2025)
How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess
von: Dionisopoulos, Lucas, et al.
Veröffentlicht: (2026)
von: Dionisopoulos, Lucas, et al.
Veröffentlicht: (2026)
Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
von: Park, Cheonbok, et al.
Veröffentlicht: (2025)
von: Park, Cheonbok, et al.
Veröffentlicht: (2025)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
von: Lee, Seanie, et al.
Veröffentlicht: (2026)
von: Lee, Seanie, et al.
Veröffentlicht: (2026)
Language over Content: Tracing Cultural Understanding in Multilingual Large Language Models
von: Cho, Seungho, et al.
Veröffentlicht: (2025)
von: Cho, Seungho, et al.
Veröffentlicht: (2025)
FIRE: Frobenius-Isometry Reinitialization for Balancing the Stability-Plasticity Tradeoff
von: Han, Isaac, et al.
Veröffentlicht: (2026)
von: Han, Isaac, et al.
Veröffentlicht: (2026)
Expressive Power of ReLU and Step Networks under Floating-Point Operations
von: Park, Yeachan, et al.
Veröffentlicht: (2024)
von: Park, Yeachan, et al.
Veröffentlicht: (2024)
Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation
von: Kim, Jaechang, et al.
Veröffentlicht: (2024)
von: Kim, Jaechang, et al.
Veröffentlicht: (2024)
Mixture of Masters: Sparse Chess Language Models with Player Routing
von: Frisoni, Giacomo, et al.
Veröffentlicht: (2026)
von: Frisoni, Giacomo, et al.
Veröffentlicht: (2026)
Can Large Language Models Reason and Optimize Under Constraints?
von: Bernier, Fabien, et al.
Veröffentlicht: (2026)
von: Bernier, Fabien, et al.
Veröffentlicht: (2026)
Hyperspherical Normalization for Scalable Deep Reinforcement Learning
von: Lee, Hojoon, et al.
Veröffentlicht: (2025)
von: Lee, Hojoon, et al.
Veröffentlicht: (2025)
Model-based Offline Reinforcement Learning with Lower Expectile Q-Learning
von: Park, Kwanyoung, et al.
Veröffentlicht: (2024)
von: Park, Kwanyoung, et al.
Veröffentlicht: (2024)
Can Large Language Models Reason and Plan?
von: Kambhampati, Subbarao
Veröffentlicht: (2024)
von: Kambhampati, Subbarao
Veröffentlicht: (2024)
Enhancing Chess Reinforcement Learning with Graph Representation
von: Rigaux, Tomas, et al.
Veröffentlicht: (2024)
von: Rigaux, Tomas, et al.
Veröffentlicht: (2024)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
von: Qu, Yun, et al.
Veröffentlicht: (2026)
von: Qu, Yun, et al.
Veröffentlicht: (2026)
Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models
von: Tao, Yongding, et al.
Veröffentlicht: (2025)
von: Tao, Yongding, et al.
Veröffentlicht: (2025)
Development and Validation of Heparin Dosing Policies Using an Offline Reinforcement Learning Algorithm
von: Lim, Yooseok, et al.
Veröffentlicht: (2024)
von: Lim, Yooseok, et al.
Veröffentlicht: (2024)
Intrinsic Task Symmetry Drives Generalization in Algorithmic Tasks
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2026)
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2026)
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
von: Park, Seonghyeon, et al.
Veröffentlicht: (2026)
von: Park, Seonghyeon, et al.
Veröffentlicht: (2026)
Large Language Models Are Better Logical Fallacy Reasoners with Counterargument, Explanation, and Goal-Aware Prompt Formulation
von: Jeong, Jiwon, et al.
Veröffentlicht: (2025)
von: Jeong, Jiwon, et al.
Veröffentlicht: (2025)
CoTox: Chain-of-Thought-Based Molecular Toxicity Reasoning and Prediction
von: Park, Jueon, et al.
Veröffentlicht: (2025)
von: Park, Jueon, et al.
Veröffentlicht: (2025)
Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning
von: Park, Jongchan, et al.
Veröffentlicht: (2025)
von: Park, Jongchan, et al.
Veröffentlicht: (2025)
TAROT: Towards Essentially Domain-Invariant Robustness with Theoretical Justification
von: Yang, Dongyoon, et al.
Veröffentlicht: (2025)
von: Yang, Dongyoon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
von: Kim, Hyunseung, et al.
Veröffentlicht: (2024) -
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
von: Kim, Donghu, et al.
Veröffentlicht: (2024) -
When Model Meets New Normals: Test-time Adaptation for Unsupervised Time-series Anomaly Detection
von: Kim, Dongmin, et al.
Veröffentlicht: (2023) -
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025) -
SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning
von: Lee, Hojoon, et al.
Veröffentlicht: (2024)