Tracking World States with Language Models: State-Based Evaluation Using Chess
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Harang, Romain, Naradowsky, Jason, Gujju, Yaswitha, Miyao, Yusuke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-Guided Ansätze Design for Quantum Circuit Born Machines in Financial Generative Modeling
von: Gujju, Yaswitha, et al.
Veröffentlicht: (2025)
von: Gujju, Yaswitha, et al.
Veröffentlicht: (2025)
Does it Chug? Towards a Data-Driven Understanding of Guitar Tone Description
von: Sutar, Pratik, et al.
Veröffentlicht: (2024)
von: Sutar, Pratik, et al.
Veröffentlicht: (2024)
How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control
von: Li, Kunhang, et al.
Veröffentlicht: (2025)
von: Li, Kunhang, et al.
Veröffentlicht: (2025)
Self-Emotion Blended Dialogue Generation in Social Simulation Agents
von: Zhang, Qiang, et al.
Veröffentlicht: (2024)
von: Zhang, Qiang, et al.
Veröffentlicht: (2024)
QuProFS: An Evolutionary Training-free Approach to Efficient Quantum Feature Map Search
von: Gujju, Yaswitha, et al.
Veröffentlicht: (2025)
von: Gujju, Yaswitha, et al.
Veröffentlicht: (2025)
Quantum Machine Learning on Near-Term Quantum Devices: Current State of Supervised and Unsupervised Techniques for Real-World Applications
von: Gujju, Yaswitha, et al.
Veröffentlicht: (2023)
von: Gujju, Yaswitha, et al.
Veröffentlicht: (2023)
ChessQA: Evaluating Large Language Models for Chess Understanding
von: Wen, Qianfeng, et al.
Veröffentlicht: (2025)
von: Wen, Qianfeng, et al.
Veröffentlicht: (2025)
A Multi-Perspective Analysis of Memorization in Large Language Models
von: Chen, Bowen, et al.
Veröffentlicht: (2024)
von: Chen, Bowen, et al.
Veröffentlicht: (2024)
A Statistical and Multi-Perspective Revisiting of the Membership Inference Attack in Large Language Models
von: Chen, Bowen, et al.
Veröffentlicht: (2024)
von: Chen, Bowen, et al.
Veröffentlicht: (2024)
Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models
von: DiSorbo, Matthew DosSantos, et al.
Veröffentlicht: (2026)
von: DiSorbo, Matthew DosSantos, et al.
Veröffentlicht: (2026)
Textless Dependency Parsing by Labeled Sequence Prediction
von: Kando, Shunsuke, et al.
Veröffentlicht: (2024)
von: Kando, Shunsuke, et al.
Veröffentlicht: (2024)
ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
Predicting Human Chess Moves: An AI Assisted Analysis of Chess Games Using Skill-group Specific n-gram Language Models
von: Zhong, Daren, et al.
Veröffentlicht: (2025)
von: Zhong, Daren, et al.
Veröffentlicht: (2025)
Beyond Accuracy: A Geometric Stability Analysis of Large Language Models in Chess Evaluation
von: Song, Xidan, et al.
Veröffentlicht: (2025)
von: Song, Xidan, et al.
Veröffentlicht: (2025)
Tracking vs. Deciding: The Dual-Capability Bottleneck in Searchless Chess Transformers
von: Li, Quanhao, et al.
Veröffentlicht: (2026)
von: Li, Quanhao, et al.
Veröffentlicht: (2026)
Grounded Chess Reasoning in Language Models via Master Distillation
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
Collaborating with AI Agents: Field Experiments on Teamwork, Productivity, and Performance
von: Ju, Harang, et al.
Veröffentlicht: (2025)
von: Ju, Harang, et al.
Veröffentlicht: (2025)
Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
von: Martorell, Nicolas, et al.
Veröffentlicht: (2026)
von: Martorell, Nicolas, et al.
Veröffentlicht: (2026)
Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models
von: Tang, Ethan
Veröffentlicht: (2026)
von: Tang, Ethan
Veröffentlicht: (2026)
Mixture of Masters: Sparse Chess Language Models with Player Routing
von: Frisoni, Giacomo, et al.
Veröffentlicht: (2026)
von: Frisoni, Giacomo, et al.
Veröffentlicht: (2026)
Toward Modeling Player-Specific Chess Behaviors
von: Sogliuzzo, Loris, et al.
Veröffentlicht: (2026)
von: Sogliuzzo, Loris, et al.
Veröffentlicht: (2026)
(How) Do Language Models Track State?
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Hybrid Dialogue State Tracking for Persian Chatbots: A Language Model-Based Approach
von: Aghabagher, Samin Mahdipour, et al.
Veröffentlicht: (2025)
von: Aghabagher, Samin Mahdipour, et al.
Veröffentlicht: (2025)
FinGen: A Dataset for Argument Generation in Finance
von: Chen, Chung-Chi, et al.
Veröffentlicht: (2024)
von: Chen, Chung-Chi, et al.
Veröffentlicht: (2024)
Do Language Models Track Entities Across State Changes?
von: Tang, Zilu, et al.
Veröffentlicht: (2026)
von: Tang, Zilu, et al.
Veröffentlicht: (2026)
Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation
von: Kim, Jaechang, et al.
Veröffentlicht: (2024)
von: Kim, Jaechang, et al.
Veröffentlicht: (2024)
LLM-State: Open World State Representation for Long-horizon Task Planning with Large Language Model
von: Chen, Siwei, et al.
Veröffentlicht: (2023)
von: Chen, Siwei, et al.
Veröffentlicht: (2023)
Evaluating In Silico Creativity: An Expert Review of AI Chess Compositions
von: Veeriah, Vivek, et al.
Veröffentlicht: (2025)
von: Veeriah, Vivek, et al.
Veröffentlicht: (2025)
Learning to Imitate with Less: Efficient Individual Behavior Modeling in Chess
von: Tang, Zhenwei, et al.
Veröffentlicht: (2025)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2025)
Aligning Large Language Models with Procedural Rules: An Autoregressive State-Tracking Prompting for In-Game Trading
von: Kim, Minkyung, et al.
Veröffentlicht: (2025)
von: Kim, Minkyung, et al.
Veröffentlicht: (2025)
Multimodal Hidden Markov Models for Persistent Emotional State Tracking
von: Ragu, Anamika, et al.
Veröffentlicht: (2026)
von: Ragu, Anamika, et al.
Veröffentlicht: (2026)
Abstract Concept Modelling in Conceptual Spaces: A Study on Chess Strategies
von: Banaee, Hadi, et al.
Veröffentlicht: (2026)
von: Banaee, Hadi, et al.
Veröffentlicht: (2026)
Maia-2: A Unified Model for Human-AI Alignment in Chess
von: Tang, Zhenwei, et al.
Veröffentlicht: (2024)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2024)
Confidence Estimation for LLM-Based Dialogue State Tracking
von: Sun, Yi-Jyun, et al.
Veröffentlicht: (2024)
von: Sun, Yi-Jyun, et al.
Veröffentlicht: (2024)
Generating Creative Chess Puzzles
von: Feng, Xidong, et al.
Veröffentlicht: (2025)
von: Feng, Xidong, et al.
Veröffentlicht: (2025)
Complete Chess Games Enable LLM Become A Chess Master
von: Zhang, Yinqi, et al.
Veröffentlicht: (2025)
von: Zhang, Yinqi, et al.
Veröffentlicht: (2025)
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2025)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2025)
Predicting User Perception of Move Brilliance in Chess
von: Zaidi, Kamron, et al.
Veröffentlicht: (2024)
von: Zaidi, Kamron, et al.
Veröffentlicht: (2024)
AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
von: Takagi, Yusuke, et al.
Veröffentlicht: (2026)
von: Takagi, Yusuke, et al.
Veröffentlicht: (2026)
Real-World Cooking Robot System from Recipes Based on Food State Recognition Using Foundation Models and PDDL
von: Kanazawa, Naoaki, et al.
Veröffentlicht: (2024)
von: Kanazawa, Naoaki, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLM-Guided Ansätze Design for Quantum Circuit Born Machines in Financial Generative Modeling
von: Gujju, Yaswitha, et al.
Veröffentlicht: (2025) -
Does it Chug? Towards a Data-Driven Understanding of Guitar Tone Description
von: Sutar, Pratik, et al.
Veröffentlicht: (2024) -
How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control
von: Li, Kunhang, et al.
Veröffentlicht: (2025) -
Self-Emotion Blended Dialogue Generation in Social Simulation Agents
von: Zhang, Qiang, et al.
Veröffentlicht: (2024) -
QuProFS: An Evolutionary Training-free Approach to Efficient Quantum Feature Map Search
von: Gujju, Yaswitha, et al.
Veröffentlicht: (2025)