Saved in:
Bibliographic Details
Main Authors: van Breda, Arco, Acar, Erman
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.03506
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915770625163264
author van Breda, Arco
Acar, Erman
author_facet van Breda, Arco
Acar, Erman
contents Following their success across many domains, transformers have also proven effective for symbolic regression (SR); however, the internal mechanisms underlying their generation of mathematical operators remain largely unexplored. Although mechanistic interpretability has successfully identified circuits in language and vision models, it has not yet been applied to SR. In this article, we introduce PATCHES, an evolutionary circuit discovery algorithm that identifies compact and correct circuits for SR. Using PATCHES, we isolate 28 circuits, providing the first circuit-level characterisation of an SR transformer. We validate these findings through a robust causal evaluation framework based on key notions such as faithfulness, completeness, and minimality. Our analysis shows that mean patching with performance-based evaluation most reliably isolates functionally correct circuits. In contrast, we demonstrate that direct logit attribution and probing classifiers primarily capture correlational features rather than causal ones, limiting their utility for circuit discovery. Overall, these results establish SR as a high-potential application domain for mechanistic interpretability and propose a principled methodology for circuit discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03506
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Explaining the Explainer: Understanding the Inner Workings of Transformer-based Symbolic Regression Models
van Breda, Arco
Acar, Erman
Machine Learning
Artificial Intelligence
Following their success across many domains, transformers have also proven effective for symbolic regression (SR); however, the internal mechanisms underlying their generation of mathematical operators remain largely unexplored. Although mechanistic interpretability has successfully identified circuits in language and vision models, it has not yet been applied to SR. In this article, we introduce PATCHES, an evolutionary circuit discovery algorithm that identifies compact and correct circuits for SR. Using PATCHES, we isolate 28 circuits, providing the first circuit-level characterisation of an SR transformer. We validate these findings through a robust causal evaluation framework based on key notions such as faithfulness, completeness, and minimality. Our analysis shows that mean patching with performance-based evaluation most reliably isolates functionally correct circuits. In contrast, we demonstrate that direct logit attribution and probing classifiers primarily capture correlational features rather than causal ones, limiting their utility for circuit discovery. Overall, these results establish SR as a high-potential application domain for mechanistic interpretability and propose a principled methodology for circuit discovery.
title Explaining the Explainer: Understanding the Inner Workings of Transformer-based Symbolic Regression Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.03506