Saved in:
| Main Authors: | Liu, Andy Zeyi, Paquette, Elliot, Sous, John |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.05683 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
by: Liu, Andy Zeyi, et al.
Published: (2026)
by: Liu, Andy Zeyi, et al.
Published: (2026)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling
by: Benigni, Lucas, et al.
Published: (2025)
by: Benigni, Lucas, et al.
Published: (2025)
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
by: Everett, Katie, et al.
Published: (2026)
by: Everett, Katie, et al.
Published: (2026)
High-Dimensional Privacy-Utility Dynamics of Noisy Stochastic Gradient Descent on Least Squares
by: Lin, Shurong, et al.
Published: (2025)
by: Lin, Shurong, et al.
Published: (2025)
4+3 Phases of Compute-Optimal Neural Scaling Laws
by: Paquette, Elliot, et al.
Published: (2024)
by: Paquette, Elliot, et al.
Published: (2024)
Logarithmic-time Schedules for Scaling Language Models with Momentum
by: Ferbach, Damien, et al.
Published: (2026)
by: Ferbach, Damien, et al.
Published: (2026)
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025)
by: Ferbach, Damien, et al.
Published: (2025)
Power-Law Spectrum of the Random Feature Model
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
Assign and Add: A Mechanistic Study of Compositional Arithmetic
by: Exoo, Brady, et al.
Published: (2026)
by: Exoo, Brady, et al.
Published: (2026)
Asymmetric Scaling Laws from Sparse Features
by: Sous, John, et al.
Published: (2026)
by: Sous, John, et al.
Published: (2026)
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
by: Xiao, Ke Liang, et al.
Published: (2024)
by: Xiao, Ke Liang, et al.
Published: (2024)
Anisotropic local law for non-separable sample covariance matrices
by: Fan, Zhou, et al.
Published: (2026)
by: Fan, Zhou, et al.
Published: (2026)
Fitting an ellipsoid to a quadratic number of random points
by: Bandeira, Afonso S., et al.
Published: (2023)
by: Bandeira, Afonso S., et al.
Published: (2023)
The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)
Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy
by: Huang, Zhendong, et al.
Published: (2026)
by: Huang, Zhendong, et al.
Published: (2026)
(Im)possibility of Automated Hallucination Detection in Large Language Models
by: Karbasi, Amin, et al.
Published: (2025)
by: Karbasi, Amin, et al.
Published: (2025)
Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra
by: Southworth, Ben S., et al.
Published: (2026)
by: Southworth, Ben S., et al.
Published: (2026)
SpectraKAN: Conditioning Spectral Operators
by: Cheng, Chun-Wun, et al.
Published: (2026)
by: Cheng, Chun-Wun, et al.
Published: (2026)
IMPROVE: Iterative Model Pipeline Refinement and Optimization Leveraging LLM Experts
by: Xue, Eric, et al.
Published: (2025)
by: Xue, Eric, et al.
Published: (2025)
LossLens: Diagnostics for Machine Learning through Loss Landscape Visual Analytics
by: Xie, Tiankai, et al.
Published: (2024)
by: Xie, Tiankai, et al.
Published: (2024)
Balanced Data, Imbalanced Spectra: Unveiling Class Disparities with Spectral Imbalance
by: Kaushik, Chiraag, et al.
Published: (2024)
by: Kaushik, Chiraag, et al.
Published: (2024)
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
by: Li, Muyang, et al.
Published: (2026)
by: Li, Muyang, et al.
Published: (2026)
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
by: Chen, Tianhao, et al.
Published: (2025)
by: Chen, Tianhao, et al.
Published: (2025)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
by: Atanasov, Alexander, et al.
Published: (2025)
by: Atanasov, Alexander, et al.
Published: (2025)
EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
by: Song, Zhiye, et al.
Published: (2026)
by: Song, Zhiye, et al.
Published: (2026)
Correctness-Optimized Residual Activation Lens (CORAL): Transferrable and Calibration-Aware Inference-Time Steering
by: Miao, Miranda Muqing, et al.
Published: (2026)
by: Miao, Miranda Muqing, et al.
Published: (2026)
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning
by: Lu, Binghang, et al.
Published: (2026)
by: Lu, Binghang, et al.
Published: (2026)
A constraints-based approach to fully interpretable neural networks for detecting learner behaviors
by: Pinto, Juan D., et al.
Published: (2025)
by: Pinto, Juan D., et al.
Published: (2025)
Towards a Unified Framework for Evaluating Explanations
by: Pinto, Juan D., et al.
Published: (2024)
by: Pinto, Juan D., et al.
Published: (2024)
Learning Physical Simulation with Message Passing Transformer
by: Xu, Zeyi, et al.
Published: (2024)
by: Xu, Zeyi, et al.
Published: (2024)
Distributional Spectral Diagnostics for Localizing Grokking Transitions
by: Wang, Ziyue, et al.
Published: (2026)
by: Wang, Ziyue, et al.
Published: (2026)
Out of the Ordinary: Spectrally Adapting Regression for Covariate Shift
by: Eyre, Benjamin, et al.
Published: (2023)
by: Eyre, Benjamin, et al.
Published: (2023)
Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps
by: Cheng, Jiashun, et al.
Published: (2025)
by: Cheng, Jiashun, et al.
Published: (2025)
Deep Learning for Educational Data Science
by: Pinto, Juan D., et al.
Published: (2024)
by: Pinto, Juan D., et al.
Published: (2024)
Asymmetric Adaptation-based Real-time Fault Diagnosis Under Transitional Operating Conditions
by: Zhao, Hongshuo, et al.
Published: (2026)
by: Zhao, Hongshuo, et al.
Published: (2026)
Performance-bounded Online Ensemble Learning Method Based on Multi-armed bandits and Its Applications in Real-time Safety Assessment
by: Hu, Songqiao, et al.
Published: (2025)
by: Hu, Songqiao, et al.
Published: (2025)
Lite-RVFL: A Lightweight Random Vector Functional-Link Neural Network for Learning Under Concept Drift
by: Hu, Songqiao, et al.
Published: (2025)
by: Hu, Songqiao, et al.
Published: (2025)
Deep Learning for Optical Misalignment Diagnostics in Multi-Lens Imaging Systems
by: Slor, Tomer, et al.
Published: (2025)
by: Slor, Tomer, et al.
Published: (2025)
Similar Items
-
Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
by: Liu, Andy Zeyi, et al.
Published: (2026) -
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024) -
Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling
by: Benigni, Lucas, et al.
Published: (2025) -
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
by: Everett, Katie, et al.
Published: (2026) -
High-Dimensional Privacy-Utility Dynamics of Noisy Stochastic Gradient Descent on Least Squares
by: Lin, Shurong, et al.
Published: (2025)