Iterative Inference in a Chess-Playing Neural Network

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sandmann, Elias, Lapuschkin, Sebastian, Samek, Wojciech
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917103103115264
author Sandmann, Elias
Lapuschkin, Sebastian
Samek, Wojciech
author_facet Sandmann, Elias
Lapuschkin, Sebastian
Samek, Wojciech
contents Do neural networks build their representations through smooth, gradual refinement, or via more complex computational processes? We investigate this by extending the logit lens to analyze the policy network of Leela Chess Zero, a superhuman chess engine. Although playing strength and puzzle-solving ability improve consistently across layers, capability progression occurs in distinct computational phases with move preferences undergoing continuous reevaluation--move rankings remain poorly correlated with final outputs until late, and correct puzzle solutions found in middle layers are sometimes overridden. This late-layer reversal is accompanied by concept preference analyses showing final layers prioritize safety over aggression, suggesting a mechanism by which heuristic priors can override tactical solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2508_21380
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Iterative Inference in a Chess-Playing Neural Network
Sandmann, Elias
Lapuschkin, Sebastian
Samek, Wojciech
Machine Learning
Artificial Intelligence
Do neural networks build their representations through smooth, gradual refinement, or via more complex computational processes? We investigate this by extending the logit lens to analyze the policy network of Leela Chess Zero, a superhuman chess engine. Although playing strength and puzzle-solving ability improve consistently across layers, capability progression occurs in distinct computational phases with move preferences undergoing continuous reevaluation--move rankings remain poorly correlated with final outputs until late, and correct puzzle solutions found in middle layers are sometimes overridden. This late-layer reversal is accompanied by concept preference analyses showing final layers prioritize safety over aggression, suggesting a mechanism by which heuristic priors can override tactical solutions.
title Iterative Inference in a Chess-Playing Neural Network
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.21380