Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Lutz, Patrick, Haris, Themistoklis, Chandra, Arjun, Gangrade, Aditya, Saligrama, Venkatesh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Linear Transformers Implicitly Discover Unified Numerical Algorithms
by: Lutz, Patrick, et al.
Published: (2025)
by: Lutz, Patrick, et al.
Published: (2025)
Constrained Linear Thompson Sampling
by: Gangrade, Aditya, et al.
Published: (2025)
by: Gangrade, Aditya, et al.
Published: (2025)
Safe Linear Bandits over Unknown Polytopes
by: Gangrade, Aditya, et al.
Published: (2022)
by: Gangrade, Aditya, et al.
Published: (2022)
Testing the Feasibility of Linear Programs with Bandit Feedback
by: Gangrade, Aditya, et al.
Published: (2024)
by: Gangrade, Aditya, et al.
Published: (2024)
Noise Stability of Transformer Models
by: Haris, Themistoklis, et al.
Published: (2026)
by: Haris, Themistoklis, et al.
Published: (2026)
Data Deletion Can Help in Adaptive RL
by: Budhraja, Param, et al.
Published: (2026)
by: Budhraja, Param, et al.
Published: (2026)
Compression Barriers for Autoregressive Transformers
by: Haris, Themistoklis, et al.
Published: (2025)
by: Haris, Themistoklis, et al.
Published: (2025)
From Compression to Expression: A Layerwise Analysis of In-Context Learning
by: Jiang, Jiachen, et al.
Published: (2025)
by: Jiang, Jiachen, et al.
Published: (2025)
$k$NN Attention Demystified: A Theoretical Exploration for Scalable Transformers
by: Haris, Themistoklis
Published: (2024)
by: Haris, Themistoklis
Published: (2024)
How Transformers Learn In-Context Recall Tasks? Optimality, Training Dynamics and Generalization
by: Nguyen, Quan, et al.
Published: (2025)
by: Nguyen, Quan, et al.
Published: (2025)
Layerwise LQR for Geometry-Aware Optimization of Deep Networks
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
by: Fan, Zhiyuan, et al.
Published: (2025)
by: Fan, Zhiyuan, et al.
Published: (2025)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
by: Ye, Chenlu, et al.
Published: (2026)
by: Ye, Chenlu, et al.
Published: (2026)
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
by: Wang, Shengao, et al.
Published: (2025)
by: Wang, Shengao, et al.
Published: (2025)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
by: Lei, Ge, et al.
Published: (2025)
by: Lei, Ge, et al.
Published: (2025)
Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs
by: Gao, Chang, et al.
Published: (2025)
by: Gao, Chang, et al.
Published: (2025)
Employing Layerwised Unsupervised Learning to Lessen Data and Loss Requirements in Forward-Forward Algorithms
by: Hwang, Taewook, et al.
Published: (2024)
by: Hwang, Taewook, et al.
Published: (2024)
Outlier-weighed Layerwise Sampling for LLM Fine-tuning
by: Li, Pengxiang, et al.
Published: (2024)
by: Li, Pengxiang, et al.
Published: (2024)
ART: Adaptive Resampling-based Training for Imbalanced Classification
by: Basandrai, Arjun, et al.
Published: (2025)
by: Basandrai, Arjun, et al.
Published: (2025)
On Some Tunable Multi-fidelity Bayesian Optimization Frameworks
by: Manoj, Arjun, et al.
Published: (2025)
by: Manoj, Arjun, et al.
Published: (2025)
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
by: He, Di, et al.
Published: (2026)
by: He, Di, et al.
Published: (2026)
In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks
by: Goel, Ayush, et al.
Published: (2026)
by: Goel, Ayush, et al.
Published: (2026)
On the Robustness of Transformers against Context Hijacking for Linear Classification
by: Li, Tianle, et al.
Published: (2025)
by: Li, Tianle, et al.
Published: (2025)
Exploring Layerwise Adversarial Robustness Through the Lens of t-SNE
by: Valentim, Inês, et al.
Published: (2024)
by: Valentim, Inês, et al.
Published: (2024)
Is Monotonic Sampling Necessary in Diffusion Models?
by: Khan, Muhammad Haris
Published: (2026)
by: Khan, Muhammad Haris
Published: (2026)
Symmetry-Aware Transformer Training for Automated Planning
by: Fritzsche, Markus, et al.
Published: (2025)
by: Fritzsche, Markus, et al.
Published: (2025)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
by: Bigelow, Eric, et al.
Published: (2025)
by: Bigelow, Eric, et al.
Published: (2025)
Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data
by: Akram, Usman, et al.
Published: (2025)
by: Akram, Usman, et al.
Published: (2025)
NANOZK: Layerwise Zero-Knowledge Proofs for Verifiable Large Language Model Inference
by: Wang, Zhaohui Geoffrey
Published: (2026)
by: Wang, Zhaohui Geoffrey
Published: (2026)
The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws
by: Zaher, Eslam, et al.
Published: (2026)
by: Zaher, Eslam, et al.
Published: (2026)
Deep Companion Learning: Enhancing Generalization Through Historical Consistency
by: Zhu, Ruizhao, et al.
Published: (2024)
by: Zhu, Ruizhao, et al.
Published: (2024)
How DNNs break the Curse of Dimensionality: Compositionality and Symmetry Learning
by: Jacot, Arthur, et al.
Published: (2024)
by: Jacot, Arthur, et al.
Published: (2024)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
by: Deora, Puneesh, et al.
Published: (2025)
by: Deora, Puneesh, et al.
Published: (2025)
Spectral Discovery of Continuous Symmetries via Generalized Fourier Transforms
by: Karjol, Pavan, et al.
Published: (2026)
by: Karjol, Pavan, et al.
Published: (2026)
Large Language Models for Imbalanced Classification: Diversity makes the difference
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Layerwise Change of Knowledge in Neural Networks
by: Cheng, Xu, et al.
Published: (2024)
by: Cheng, Xu, et al.
Published: (2024)
Investigation into In-Context Learning Capabilities of Transformers
by: Chandrupatla, Rushil, et al.
Published: (2026)
by: Chandrupatla, Rushil, et al.
Published: (2026)
Orion-Bix: Bi-Axial Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025)
by: Bouadi, Mohamed, et al.
Published: (2025)
A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
by: Song, Bingqing, et al.
Published: (2025)
by: Song, Bingqing, et al.
Published: (2025)
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025)
by: Bouadi, Mohamed, et al.
Published: (2025)
Similar Items
-
Linear Transformers Implicitly Discover Unified Numerical Algorithms
by: Lutz, Patrick, et al.
Published: (2025) -
Constrained Linear Thompson Sampling
by: Gangrade, Aditya, et al.
Published: (2025) -
Safe Linear Bandits over Unknown Polytopes
by: Gangrade, Aditya, et al.
Published: (2022) -
Testing the Feasibility of Linear Programs with Bandit Feedback
by: Gangrade, Aditya, et al.
Published: (2024) -
Noise Stability of Transformer Models
by: Haris, Themistoklis, et al.
Published: (2026)