Saved in:
| Main Authors: | Xie, Shuo, Wang, Tianhao, Reddi, Sashank, Kumar, Sanjiv, Li, Zhiyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.10537 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning with Latent Thoughts: On the Power of Looped Transformers
by: Saunshi, Nikunj, et al.
Published: (2025)
by: Saunshi, Nikunj, et al.
Published: (2025)
Landscape-Aware Growing: The Power of a Little LAG
by: Karp, Stefani, et al.
Published: (2024)
by: Karp, Stefani, et al.
Published: (2024)
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
by: Xie, Shuo, et al.
Published: (2025)
by: Xie, Shuo, et al.
Published: (2025)
Simplicity Bias via Global Convergence of Sharpness Minimization
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
On the Inductive Bias of Stacking Towards Improving Reasoning
by: Saunshi, Nikunj, et al.
Published: (2024)
by: Saunshi, Nikunj, et al.
Published: (2024)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
by: Yadav, Robin, et al.
Published: (2025)
by: Yadav, Robin, et al.
Published: (2025)
Efficient Document Ranking with Learnable Late Interactions
by: Ji, Ziwei, et al.
Published: (2024)
by: Ji, Ziwei, et al.
Published: (2024)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Improving Adaptive Moment Optimization via Preconditioner Diagonalization
by: Nguyen, Son, et al.
Published: (2025)
by: Nguyen, Son, et al.
Published: (2025)
Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation
by: Lukasik, Michal, et al.
Published: (2025)
by: Lukasik, Michal, et al.
Published: (2025)
Adaptive Preconditioners Trigger Loss Spikes in Adam
by: Bai, Zhiwei, et al.
Published: (2025)
by: Bai, Zhiwei, et al.
Published: (2025)
Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Unifying Sequences, Structures, and Descriptions for Any-to-Any Protein Generation with the Large Multimodal Model HelixProtX
by: Chen, Zhiyuan, et al.
Published: (2024)
by: Chen, Zhiyuan, et al.
Published: (2024)
Deep Learning Agents Trained For Avoidance Behave Like Hawks And Doves
by: Reddi, Aryaman
Published: (2025)
by: Reddi, Aryaman
Published: (2025)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
by: Rawat, Ankit Singh, et al.
Published: (2024)
by: Rawat, Ankit Singh, et al.
Published: (2024)
Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
by: Mohamadi, Mohamad Amin, et al.
Published: (2025)
by: Mohamadi, Mohamad Amin, et al.
Published: (2025)
TinyTorch: Building Machine Learning Systems from First Principles
by: Reddi, Vijay Janapa
Published: (2026)
by: Reddi, Vijay Janapa
Published: (2026)
Generative modeling of Sparse Approximate Inverse Preconditioners
by: Li, Mou, et al.
Published: (2024)
by: Li, Mou, et al.
Published: (2024)
Are More Tokens Rational? Inference-Time Scaling in Language Models as Adaptive Resource Rationality
by: Hu, Zhimin, et al.
Published: (2026)
by: Hu, Zhimin, et al.
Published: (2026)
Investigation of Compressor Cascade Flow Using Physics- Informed Neural Networks with Adaptive Learning Strategy
by: Li, Zhihui, et al.
Published: (2023)
by: Li, Zhihui, et al.
Published: (2023)
A Non-asymptotic Analysis for Learning and Applying a Preconditioner in MCMC
by: Hird, Max, et al.
Published: (2026)
by: Hird, Max, et al.
Published: (2026)
Reasoning Distillation for Lightweight Automated Program Repair
by: Balasubramanian, Aanand, et al.
Published: (2026)
by: Balasubramanian, Aanand, et al.
Published: (2026)
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)
by: Morwani, Depen, et al.
Published: (2024)
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
by: Kwon, Soo Min, et al.
Published: (2026)
by: Kwon, Soo Min, et al.
Published: (2026)
Understanding the Countably Infinite: Neural Network Models of the Successor Function and its Acquisition
by: Gupta, Vima, et al.
Published: (2023)
by: Gupta, Vima, et al.
Published: (2023)
Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
Generative AI Agents in Autonomous Machines: A Safety Perspective
by: Jabbour, Jason, et al.
Published: (2024)
by: Jabbour, Jason, et al.
Published: (2024)
LAuReL: Learned Augmented Residual Layer
by: Menghani, Gaurav, et al.
Published: (2024)
by: Menghani, Gaurav, et al.
Published: (2024)
Unified Convergence Analysis for Adaptive Optimization with Moving Average Estimator
by: Guo, Zhishuai, et al.
Published: (2021)
by: Guo, Zhishuai, et al.
Published: (2021)
Curvature-Informed SGD via General Purpose Lie-Group Preconditioners
by: Pooladzandi, Omead, et al.
Published: (2024)
by: Pooladzandi, Omead, et al.
Published: (2024)
Enhanced High-Dimensional Data Visualization through Adaptive Multi-Scale Manifold Embedding
by: Ni, Tianhao, et al.
Published: (2025)
by: Ni, Tianhao, et al.
Published: (2025)
Optimizing Few-Step Generation with Adaptive Matching Distillation
by: Bai, Lichen, et al.
Published: (2026)
by: Bai, Lichen, et al.
Published: (2026)
Unified Transfer Learning Models in High-Dimensional Linear Regression
by: Liu, Shuo Shuo
Published: (2023)
by: Liu, Shuo Shuo
Published: (2023)
FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
by: Li, Jiaoyang, et al.
Published: (2025)
by: Li, Jiaoyang, et al.
Published: (2025)
Relative Kinetic Utility for Reasoning-Aware Structural Pruning in Large Language Models
by: Qian, Tianhao
Published: (2026)
by: Qian, Tianhao
Published: (2026)
Preconditioners for the Stochastic Training of Neural Fields
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
Similar Items
-
Reasoning with Latent Thoughts: On the Power of Looped Transformers
by: Saunshi, Nikunj, et al.
Published: (2025) -
Landscape-Aware Growing: The Power of a Little LAG
by: Karp, Stefani, et al.
Published: (2024) -
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024) -
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
by: Gatmiry, Khashayar, et al.
Published: (2024) -
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
by: Xie, Shuo, et al.
Published: (2025)