Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers
Fuente:
arXiv
Saved in:
| Main Author: | Xu, Yongzhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Spectral Edge Dynamics Reveal Functional Modes of Learning
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Spectral Edge Dynamics of Training Trajectories: Signal--Noise Geometry Across Scales
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Spectral Edge Dynamics: An Analytical-Empirical Study of Phase Transitions in Neural Network Training
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
by: Bali, Karan, et al.
Published: (2026)
by: Bali, Karan, et al.
Published: (2026)
Optimizer-Induced Low-Dimensional Drift and Transverse Dynamics in Transformer Training
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Knowledge Circuits in Pretrained Transformers
by: Yao, Yunzhi, et al.
Published: (2024)
by: Yao, Yunzhi, et al.
Published: (2024)
Low-Dimensional Execution Manifolds in Transformer Learning Dynamics: Evidence from Modular Arithmetic Tasks
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Global Low-Rank, Local Full-Rank: The Holographic Encoding of Learned Algorithms
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Early-Warning Signals of Grokking via Loss-Landscape Geometry
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Gradient-Direction Sensitivity Reveals Linear-Centroid Coupling Hidden by Optimizer Trajectories
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
The Geometry of Multi-Task Grokking: Transverse Instability, Superposition, and Weight Decay Phase Structure
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Low-Dimensional and Transversely Curved Optimization Dynamics in Grokking
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Neuronal Attention Circuit (NAC) for Representation Learning
by: Razzaq, Waleed, et al.
Published: (2025)
by: Razzaq, Waleed, et al.
Published: (2025)
Circuits, Features, and Heuristics in Molecular Transformers
by: Varadi, Kristof, et al.
Published: (2025)
by: Varadi, Kristof, et al.
Published: (2025)
CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection
by: Panuganti, Rajkiran
Published: (2026)
by: Panuganti, Rajkiran
Published: (2026)
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
by: Adhikari, Rabin
Published: (2025)
by: Adhikari, Rabin
Published: (2025)
Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
by: Chrisman, Brianna, et al.
Published: (2025)
by: Chrisman, Brianna, et al.
Published: (2025)
Neuronal Stochastic Attention Circuit (NSAC) for Probabilistic Representation Learning
by: Razzaq, Waleed, et al.
Published: (2026)
by: Razzaq, Waleed, et al.
Published: (2026)
Sparse Attention Decomposition Applied to Circuit Tracing
by: Franco, Gabriel, et al.
Published: (2024)
by: Franco, Gabriel, et al.
Published: (2024)
SynCircuit: Automated Generation of New Synthetic RTL Circuits Can Enable Big Data in Circuits
by: Liu, Shang, et al.
Published: (2025)
by: Liu, Shang, et al.
Published: (2025)
GTAC: A Generative Transformer for Approximate Circuits
by: Wang, Jingxin, et al.
Published: (2025)
by: Wang, Jingxin, et al.
Published: (2025)
Transformer Circuit Faithfulness Metrics are not Robust
by: Miller, Joseph, et al.
Published: (2024)
by: Miller, Joseph, et al.
Published: (2024)
Uncovering Graph Reasoning in Decoder-only Transformers with Circuit Tracing
by: Dai, Xinnan, et al.
Published: (2025)
by: Dai, Xinnan, et al.
Published: (2025)
A Step Toward Federated Pretraining of Multimodal Large Language Models
by: Xiong, Baochen, et al.
Published: (2026)
by: Xiong, Baochen, et al.
Published: (2026)
Circuit Complexity of Hierarchical Knowledge Tracing and Implications for Log-Precision Transformers
by: Liu, Naiming, et al.
Published: (2026)
by: Liu, Naiming, et al.
Published: (2026)
Inverse Design in Distributed Circuits Using Single-Step Reinforcement Learning
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Average Attention Transformers and Arithmetic Circuits
by: Ehrmuth, Lena, et al.
Published: (2026)
by: Ehrmuth, Lena, et al.
Published: (2026)
Causal Neural Probabilistic Circuits
by: Chen, Weixin, et al.
Published: (2026)
by: Chen, Weixin, et al.
Published: (2026)
Restructuring Tractable Probabilistic Circuits
by: Zhang, Honghua, et al.
Published: (2024)
by: Zhang, Honghua, et al.
Published: (2024)
Soft Learning Probabilistic Circuits
by: Ghandi, Soroush, et al.
Published: (2024)
by: Ghandi, Soroush, et al.
Published: (2024)
Sparse Probabilistic Graph Circuits
by: Rektoris, Martin, et al.
Published: (2025)
by: Rektoris, Martin, et al.
Published: (2025)
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
by: Skelic, Lejla, et al.
Published: (2025)
by: Skelic, Lejla, et al.
Published: (2025)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
by: Patel, Dev, et al.
Published: (2025)
by: Patel, Dev, et al.
Published: (2025)
SymCircuit: Bayesian Structure Inference for Tractable Probabilistic Circuits via Entropy-Regularized Reinforcement Learning
by: Ju, Y. Sungtaek
Published: (2026)
by: Ju, Y. Sungtaek
Published: (2026)
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
by: Huang, Ruizhe, et al.
Published: (2025)
by: Huang, Ruizhe, et al.
Published: (2025)
Hypothesis Testing the Circuit Hypothesis in LLMs
by: Shi, Claudia, et al.
Published: (2024)
by: Shi, Claudia, et al.
Published: (2024)
Probabilistic Circuits for Cumulative Distribution Functions
by: Broadrick, Oliver, et al.
Published: (2024)
by: Broadrick, Oliver, et al.
Published: (2024)
Similar Items
-
When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures
by: Xu, Yongzhong
Published: (2026) -
Spectral Edge Dynamics Reveal Functional Modes of Learning
by: Xu, Yongzhong
Published: (2026) -
Spectral Edge Dynamics of Training Trajectories: Signal--Noise Geometry Across Scales
by: Xu, Yongzhong
Published: (2026) -
Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
by: Xu, Yongzhong
Published: (2026) -
Spectral Edge Dynamics: An Analytical-Empirical Study of Phase Transitions in Neural Network Training
by: Xu, Yongzhong
Published: (2026)