Optimizer-Induced Low-Dimensional Drift and Transverse Dynamics in Transformer Training
Fuente:
arXiv
Saved in:
| Main Author: | Xu, Yongzhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Low-Dimensional and Transversely Curved Optimization Dynamics in Grokking
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Low-Dimensional Execution Manifolds in Transformer Learning Dynamics: Evidence from Modular Arithmetic Tasks
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
The Geometry of Multi-Task Grokking: Transverse Instability, Superposition, and Weight Decay Phase Structure
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Spectral Edge Dynamics of Training Trajectories: Signal--Noise Geometry Across Scales
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Spectral Edge Dynamics: An Analytical-Empirical Study of Phase Transitions in Neural Network Training
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Global Low-Rank, Local Full-Rank: The Holographic Encoding of Learned Algorithms
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Spectral Edge Dynamics Reveal Functional Modes of Learning
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Gradient-Direction Sensitivity Reveals Linear-Centroid Coupling Hidden by Optimizer Trajectories
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Early-Warning Signals of Grokking via Loss-Landscape Geometry
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
G-Drift MIA: Membership Inference via Gradient-Induced Feature Drift in LLMs
by: Ranjan, Ravi, et al.
Published: (2026)
by: Ranjan, Ravi, et al.
Published: (2026)
High-Dimensional Search, Low-Dimensional Solution: Decoupling Optimization from Representation
by: Kalyoncuoglu, Yusuf, et al.
Published: (2025)
by: Kalyoncuoglu, Yusuf, et al.
Published: (2025)
Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
by: Xu, Alec S., et al.
Published: (2026)
by: Xu, Alec S., et al.
Published: (2026)
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
Optimization-Induced Dynamics of Lipschitz Continuity in Neural Networks
by: Luo, Róisín, et al.
Published: (2025)
by: Luo, Róisín, et al.
Published: (2025)
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
by: Fluri, Lukas, et al.
Published: (2024)
by: Fluri, Lukas, et al.
Published: (2024)
Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
by: Xu, Cong, et al.
Published: (2025)
by: Xu, Cong, et al.
Published: (2025)
REG: A Regularization Optimizer for Robust Training Dynamics
by: Liu, Zehua, et al.
Published: (2025)
by: Liu, Zehua, et al.
Published: (2025)
Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
by: Ren, Ruifeng, et al.
Published: (2025)
by: Ren, Ruifeng, et al.
Published: (2025)
Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
by: Yousuf, Raquib Bin, et al.
Published: (2025)
by: Yousuf, Raquib Bin, et al.
Published: (2025)
How Transformers Learn In-Context Recall Tasks? Optimality, Training Dynamics and Generalization
by: Nguyen, Quan, et al.
Published: (2025)
by: Nguyen, Quan, et al.
Published: (2025)
Zero-Training Temporal Drift Detection for Transformer Sentiment Models: A Comprehensive Analysis on Authentic Social Media Streams
by: Bansal, Aayam, et al.
Published: (2025)
by: Bansal, Aayam, et al.
Published: (2025)
Federated Learning with Sample-level Client Drift Mitigation
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm
by: Shashidhar, Sarvesh, et al.
Published: (2025)
by: Shashidhar, Sarvesh, et al.
Published: (2025)
DriftXpress: Faster Drifting Models via Projected RKHS Fields
by: Falahati, Ali, et al.
Published: (2026)
by: Falahati, Ali, et al.
Published: (2026)
Dynamic Low-rank Approximation of Full-Matrix Preconditioner for Training Generalized Linear Models
by: Matveeva, Tatyana, et al.
Published: (2025)
by: Matveeva, Tatyana, et al.
Published: (2025)
Machine Unlearning in Low-Dimensional Feature Subspace
by: Fang, Kun, et al.
Published: (2026)
by: Fang, Kun, et al.
Published: (2026)
Drift Flow Matching
by: Ma, Chenrui, et al.
Published: (2026)
by: Ma, Chenrui, et al.
Published: (2026)
Drift Q-Learning
by: Houssaini, Anas, et al.
Published: (2026)
by: Houssaini, Anas, et al.
Published: (2026)
CORAL: Concept Drift Representation Learning for Co-evolving Time-series
by: Xu, Kunpeng, et al.
Published: (2025)
by: Xu, Kunpeng, et al.
Published: (2025)
Dynamic TMoE: A Drift-Aware Dynamic Mixture of Experts Framework for Non-Stationary Time Series Forecasting
by: Zhu, Jiawen, et al.
Published: (2026)
by: Zhu, Jiawen, et al.
Published: (2026)
WormKAN: Are KAN Effective for Identifying and Tracking Concept Drift in Time Series?
by: Xu, Kunpeng, et al.
Published: (2024)
by: Xu, Kunpeng, et al.
Published: (2024)
Taming Preconditioner Drift: Unlocking the Potential of Second-Order Optimizers for Federated Learning on Non-IID Data
by: Liu, Junkang, et al.
Published: (2026)
by: Liu, Junkang, et al.
Published: (2026)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
by: Lee, Jung Hyun, et al.
Published: (2024)
by: Lee, Jung Hyun, et al.
Published: (2024)
Exploring Dynamic Properties of Backdoor Training Through Information Bottleneck
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Dynamic Influence Tracker: Measuring Time-Varying Sample Influence During Training
by: Xu, Jie, et al.
Published: (2025)
by: Xu, Jie, et al.
Published: (2025)
Error Distribution Smoothing:Advancing Low-Dimensional Imbalanced Regression
by: Chen, Donghe, et al.
Published: (2025)
by: Chen, Donghe, et al.
Published: (2025)
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
by: Modoranu, Ionut-Vlad, et al.
Published: (2025)
by: Modoranu, Ionut-Vlad, et al.
Published: (2025)
Similar Items
-
Low-Dimensional and Transversely Curved Optimization Dynamics in Grokking
by: Xu, Yongzhong
Published: (2026) -
Low-Dimensional Execution Manifolds in Transformer Learning Dynamics: Evidence from Modular Arithmetic Tasks
by: Xu, Yongzhong
Published: (2026) -
The Geometry of Multi-Task Grokking: Transverse Instability, Superposition, and Weight Decay Phase Structure
by: Xu, Yongzhong
Published: (2026) -
Spectral Edge Dynamics of Training Trajectories: Signal--Noise Geometry Across Scales
by: Xu, Yongzhong
Published: (2026) -
Spectral Edge Dynamics: An Analytical-Empirical Study of Phase Transitions in Neural Network Training
by: Xu, Yongzhong
Published: (2026)