Similar Items
Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis
by: Fartale, Harshwardhan, et al.
Published: (2025)
by: Fartale, Harshwardhan, et al.
Published: (2025)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
by: Yehudai, Gilad, et al.
Published: (2025)
by: Yehudai, Gilad, et al.
Published: (2025)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
by: Fu, Zihao, et al.
Published: (2025)
by: Fu, Zihao, et al.
Published: (2025)
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Transformer Normalisation Layers and the Independence of Semantic Subspaces
by: Menary, Stephen, et al.
Published: (2024)
by: Menary, Stephen, et al.
Published: (2024)
Dynamic Layer Tying for Parameter-Efficient Transformers
by: Hay, Tamir David, et al.
Published: (2024)
by: Hay, Tamir David, et al.
Published: (2024)
Generalizable Reasoning through Compositional Energy Minimization
by: Oarga, Alexandru, et al.
Published: (2025)
by: Oarga, Alexandru, et al.
Published: (2025)
Universal Approximation Theorem for a Single-Layer Transformer
by: Gumaan, Esmail
Published: (2025)
by: Gumaan, Esmail
Published: (2025)
WARP: Guaranteed Inner-Layer Repair of NLP Transformers
by: Hsu, Hsin-Ling, et al.
Published: (2026)
by: Hsu, Hsin-Ling, et al.
Published: (2026)
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
by: Shrestha, Safal, et al.
Published: (2026)
by: Shrestha, Safal, et al.
Published: (2026)
Can Large Language Models Reason and Optimize Under Constraints?
by: Bernier, Fabien, et al.
Published: (2026)
by: Bernier, Fabien, et al.
Published: (2026)
Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models
by: Lippl, Samuel, et al.
Published: (2025)
by: Lippl, Samuel, et al.
Published: (2025)
On the Role of Transformer Feed-Forward Layers in Nonlinear In-Context Learning
by: Sun, Haoyuan, et al.
Published: (2025)
by: Sun, Haoyuan, et al.
Published: (2025)
Engineering Verifiable Modularity in Transformers via Per-Layer Supervision
by: Kerce, J. Clayton
Published: (2026)
by: Kerce, J. Clayton
Published: (2026)
Revisiting Transformer Layer Parameterization Through Causal Energy Minimization
by: Xu, Jin, et al.
Published: (2026)
by: Xu, Jin, et al.
Published: (2026)
Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers
by: Tang, Mohan, et al.
Published: (2026)
by: Tang, Mohan, et al.
Published: (2026)
Revealing Combinatorial Reasoning of GNNs via Graph Concept Bottleneck Layer
by: Niu, Yue, et al.
Published: (2026)
by: Niu, Yue, et al.
Published: (2026)
Strassen Attention, Split VC Dimension and Compositionality in Transformers
by: Kozachinskiy, Alexander, et al.
Published: (2025)
by: Kozachinskiy, Alexander, et al.
Published: (2025)
On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery
by: Liu, Renpu, et al.
Published: (2024)
by: Liu, Renpu, et al.
Published: (2024)
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel
by: Cutler, Dylan, et al.
Published: (2025)
by: Cutler, Dylan, et al.
Published: (2025)
Multi-Layer Attention-Based Explainability via Transformers for Tabular Data
by: Gavito, Andrea Treviño, et al.
Published: (2023)
by: Gavito, Andrea Treviño, et al.
Published: (2023)
SGFormer: Single-Layer Graph Transformers with Approximation-Free Linear Complexity
by: Wu, Qitian, et al.
Published: (2024)
by: Wu, Qitian, et al.
Published: (2024)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
by: Poduval, Prathyush, et al.
Published: (2026)
by: Poduval, Prathyush, et al.
Published: (2026)
Uncovering Layer-Dependent Activation Sparsity Patterns in ReLU Transformers
by: Wild, Cody, et al.
Published: (2024)
by: Wild, Cody, et al.
Published: (2024)
OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World Data
by: Renda, Alana, et al.
Published: (2025)
by: Renda, Alana, et al.
Published: (2025)
Understanding Transformer Reasoning Capabilities via Graph Algorithms
by: Sanford, Clayton, et al.
Published: (2024)
by: Sanford, Clayton, et al.
Published: (2024)
Ordinary Least Squares is a Special Case of Transformer
by: Tan, Xiaojun, et al.
Published: (2026)
by: Tan, Xiaojun, et al.
Published: (2026)
Channel Merging: Preserving Specialization for Merged Experts
by: Zhang, Mingyang, et al.
Published: (2024)
by: Zhang, Mingyang, et al.
Published: (2024)
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
by: Nepal, Aadim, et al.
Published: (2025)
by: Nepal, Aadim, et al.
Published: (2025)
Uncovering Graph Reasoning in Decoder-only Transformers with Circuit Tracing
by: Dai, Xinnan, et al.
Published: (2025)
by: Dai, Xinnan, et al.
Published: (2025)
Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
by: Dhayalkar, Sahil Rajesh
Published: (2025)
by: Dhayalkar, Sahil Rajesh
Published: (2025)
Beyond the Known: Decision Making with Counterfactual Reasoning Decision Transformer
by: Nguyen, Minh Hoang, et al.
Published: (2025)
by: Nguyen, Minh Hoang, et al.
Published: (2025)
The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
by: Pengmei, Zihan, et al.
Published: (2025)
by: Pengmei, Zihan, et al.
Published: (2025)
Representational Homomorphism Predicts and Improves Compositional Generalization In Transformer Language Model
by: An, Zhiyu, et al.
Published: (2026)
by: An, Zhiyu, et al.
Published: (2026)
Neural Probabilistic Circuits: Enabling Compositional and Interpretable Predictions through Logical Reasoning
by: Chen, Weixin, et al.
Published: (2025)
by: Chen, Weixin, et al.
Published: (2025)
Stability of Transformers under Layer Normalization
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs
by: Wang, Jing, et al.
Published: (2026)
by: Wang, Jing, et al.
Published: (2026)
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
A Cost-Sensitive Transformer Model for Prognostics Under Highly Imbalanced Industrial Data
by: Beikmohammadi, Ali, et al.
Published: (2024)
by: Beikmohammadi, Ali, et al.
Published: (2024)
On The Specialization of Neural Modules
by: Jarvis, Devon, et al.
Published: (2024)
by: Jarvis, Devon, et al.
Published: (2024)
Similar Items
-
Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis
by: Fartale, Harshwardhan, et al.
Published: (2025) -
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
by: Yehudai, Gilad, et al.
Published: (2025) -
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
by: Fu, Zihao, et al.
Published: (2025) -
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024) -
Transformer Normalisation Layers and the Independence of Semantic Subspaces
by: Menary, Stephen, et al.
Published: (2024)