Beyond Accuracy: A Geometric Stability Analysis of Large Language Models in Chess Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Xidan, Wang, Weiqi, Cao, Ruifeng, Hu, Qingya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ChessQA: Evaluating Large Language Models for Chess Understanding
von: Wen, Qianfeng, et al.
Veröffentlicht: (2025)
von: Wen, Qianfeng, et al.
Veröffentlicht: (2025)
ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
Assisting Research Proposal Writing with Large Language Models: Evaluation and Refinement
von: Ren, Jing, et al.
Veröffentlicht: (2025)
von: Ren, Jing, et al.
Veröffentlicht: (2025)
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Tracking World States with Language Models: State-Based Evaluation Using Chess
von: Harang, Romain, et al.
Veröffentlicht: (2025)
von: Harang, Romain, et al.
Veröffentlicht: (2025)
Predicting Human Chess Moves: An AI Assisted Analysis of Chess Games Using Skill-group Specific n-gram Language Models
von: Zhong, Daren, et al.
Veröffentlicht: (2025)
von: Zhong, Daren, et al.
Veröffentlicht: (2025)
Measuring Stability Beyond Accuracy in Small Open-Source Medical Large Language Models for Pediatric Endocrinology
von: D'Amario, Vanessa, et al.
Veröffentlicht: (2025)
von: D'Amario, Vanessa, et al.
Veröffentlicht: (2025)
Grounded Chess Reasoning in Language Models via Master Distillation
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
Beyond Multiple-Choice Accuracy: Real-World Challenges of Implementing Large Language Models in Healthcare
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Personalized Large Language Model Assistant with Evolving Conditional Memory
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2023)
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2023)
Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
von: Huang, Liangjie, et al.
Veröffentlicht: (2025)
von: Huang, Liangjie, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Characterizing Code Comprehension Capabilities in (Large) Language Models
von: Mächtle, Felix, et al.
Veröffentlicht: (2026)
von: Mächtle, Felix, et al.
Veröffentlicht: (2026)
VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
von: Li, Weiqi, et al.
Veröffentlicht: (2025)
von: Li, Weiqi, et al.
Veröffentlicht: (2025)
The ORCA Benchmark: Evaluating Real-World Calculation Accuracy in Large Language Models
von: Herambourg, Claudia, et al.
Veröffentlicht: (2025)
von: Herambourg, Claudia, et al.
Veröffentlicht: (2025)
Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models
von: Tang, Ethan
Veröffentlicht: (2026)
von: Tang, Ethan
Veröffentlicht: (2026)
Mixture of Masters: Sparse Chess Language Models with Player Routing
von: Frisoni, Giacomo, et al.
Veröffentlicht: (2026)
von: Frisoni, Giacomo, et al.
Veröffentlicht: (2026)
Toward Modeling Player-Specific Chess Behaviors
von: Sogliuzzo, Loris, et al.
Veröffentlicht: (2026)
von: Sogliuzzo, Loris, et al.
Veröffentlicht: (2026)
GraphScout: Empowering Large Language Models with Intrinsic Exploration Ability for Agentic Graph Reasoning
von: Ying, Yuchen, et al.
Veröffentlicht: (2026)
von: Ying, Yuchen, et al.
Veröffentlicht: (2026)
Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks
von: Pimentel, Marco AF, et al.
Veröffentlicht: (2024)
von: Pimentel, Marco AF, et al.
Veröffentlicht: (2024)
Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning
von: Yang, Xia, et al.
Veröffentlicht: (2026)
von: Yang, Xia, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Evaluating Forecasting Models by Multi-Echelon Inventory Cost
von: Marik, Swata, et al.
Veröffentlicht: (2026)
von: Marik, Swata, et al.
Veröffentlicht: (2026)
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2025)
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2025)
Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models
von: Mouselinos, Spyridon, et al.
Veröffentlicht: (2024)
von: Mouselinos, Spyridon, et al.
Veröffentlicht: (2024)
Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation
von: Kim, Jaechang, et al.
Veröffentlicht: (2024)
von: Kim, Jaechang, et al.
Veröffentlicht: (2024)
MoZIP: A Multilingual Benchmark to Evaluate Large Language Models in Intellectual Property
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
An Information-Geometric Framework for Stability Analysis of Large Language Models under Entropic Stress
von: Karimov, Hikmat, et al.
Veröffentlicht: (2026)
von: Karimov, Hikmat, et al.
Veröffentlicht: (2026)
Abstract Concept Modelling in Conceptual Spaces: A Study on Chess Strategies
von: Banaee, Hadi, et al.
Veröffentlicht: (2026)
von: Banaee, Hadi, et al.
Veröffentlicht: (2026)
Maia-2: A Unified Model for Human-AI Alignment in Chess
von: Tang, Zhenwei, et al.
Veröffentlicht: (2024)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2024)
Complete Chess Games Enable LLM Become A Chess Master
von: Zhang, Yinqi, et al.
Veröffentlicht: (2025)
von: Zhang, Yinqi, et al.
Veröffentlicht: (2025)
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
von: Ruoss, Anian, et al.
Veröffentlicht: (2024)
von: Ruoss, Anian, et al.
Veröffentlicht: (2024)
Evaluating In Silico Creativity: An Expert Review of AI Chess Compositions
von: Veeriah, Vivek, et al.
Veröffentlicht: (2025)
von: Veeriah, Vivek, et al.
Veröffentlicht: (2025)
Learning to Imitate with Less: Efficient Individual Behavior Modeling in Chess
von: Tang, Zhenwei, et al.
Veröffentlicht: (2025)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2025)
A Comprehensive Evaluation of Large Language Models on Aspect-Based Sentiment Analysis
von: Zhou, Changzhi, et al.
Veröffentlicht: (2024)
von: Zhou, Changzhi, et al.
Veröffentlicht: (2024)
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
NUMCoT: Numerals and Units of Measurement in Chain-of-Thought Reasoning using Large Language Models
von: Xu, Ancheng, et al.
Veröffentlicht: (2024)
von: Xu, Ancheng, et al.
Veröffentlicht: (2024)
StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis
von: Song, Xinyi, et al.
Veröffentlicht: (2025)
von: Song, Xinyi, et al.
Veröffentlicht: (2025)
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
von: Zhou, Yongxi, et al.
Veröffentlicht: (2026)
von: Zhou, Yongxi, et al.
Veröffentlicht: (2026)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
von: Hao, Yijie, et al.
Veröffentlicht: (2025)
von: Hao, Yijie, et al.
Veröffentlicht: (2025)
Evaluating the Feasibility and Accuracy of Large Language Models for Medical History-Taking in Obstetrics and Gynecology
von: Liu, Dou, et al.
Veröffentlicht: (2025)
von: Liu, Dou, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ChessQA: Evaluating Large Language Models for Chess Understanding
von: Wen, Qianfeng, et al.
Veröffentlicht: (2025) -
ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
von: Liu, Jincheng, et al.
Veröffentlicht: (2025) -
Assisting Research Proposal Writing with Large Language Models: Evaluation and Refinement
von: Ren, Jing, et al.
Veröffentlicht: (2025) -
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026) -
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)