Saved in:
| Main Authors: | Lau, Tim Tsz-Kit, Su, Weijie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.18106 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
by: Yadav, Sarthak, et al.
Published: (2025)
by: Yadav, Sarthak, et al.
Published: (2025)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
by: Su, Weijie
Published: (2025)
by: Su, Weijie
Published: (2025)
Depth Registers Unlock W4A4 on SwiGLU: A Reader/Generator Decomposition
by: Liu, Ziyang
Published: (2026)
by: Liu, Ziyang
Published: (2026)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
by: Farhat, Yehya, et al.
Published: (2023)
by: Farhat, Yehya, et al.
Published: (2023)
First-Order Geometry, Spectral Compression, and Structural Compatibility under Bounded Computation
by: Li, Changkai
Published: (2026)
by: Li, Changkai
Published: (2026)
Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation
by: Jantsch, Lasse Marten, et al.
Published: (2026)
by: Jantsch, Lasse Marten, et al.
Published: (2026)
Solving Functional Optimization with Deep Networks and Variational Principles
by: Kamtue, Kawisorn, et al.
Published: (2024)
by: Kamtue, Kawisorn, et al.
Published: (2024)
Distributionally Robust Free Energy Principle for Decision-Making
by: Shafiei, Allahkaram, et al.
Published: (2025)
by: Shafiei, Allahkaram, et al.
Published: (2025)
LoLA-SpecViT: Local Attention SwiGLU Vision Transformer with LoRA for Hyperspectral Imaging
by: Zidi, Fadi Abdeladhim, et al.
Published: (2025)
by: Zidi, Fadi Abdeladhim, et al.
Published: (2025)
Let's Have a Conversation: Designing and Evaluating LLM Agents for Interactive Optimization
by: Drossman, Joshua, et al.
Published: (2026)
by: Drossman, Joshua, et al.
Published: (2026)
Accelerated Decentralized Constraint-Coupled Optimization: A Dual$^2$ Approach
by: Li, Jingwang, et al.
Published: (2025)
by: Li, Jingwang, et al.
Published: (2025)
Reward Collapse in Aligning Large Language Models
by: Song, Ziang, et al.
Published: (2023)
by: Song, Ziang, et al.
Published: (2023)
Adjoint-Compatible Surrogates of the Expected Information Gain for Optimal Experimental Design
by: de Montella, Luc, et al.
Published: (2026)
by: de Montella, Luc, et al.
Published: (2026)
The Principle of Proportional Duty: A Knowledge-Duty Framework for Ethical Equilibrium in Human and Artificial Systems
by: Prescher, Timothy
Published: (2025)
by: Prescher, Timothy
Published: (2025)
BoxLitE: A Faithful Knowledge Base Embedding Based on Convex Optimization
by: Lourenço, Bruno F., et al.
Published: (2026)
by: Lourenço, Bruno F., et al.
Published: (2026)
Hereditary Geometric Meta-RL: Nonlocal Generalization via Task Symmetries
by: Nitschke, Paul, et al.
Published: (2026)
by: Nitschke, Paul, et al.
Published: (2026)
Robust Learning Rate Selection for Stochastic Optimization via Splitting Diagnostic
by: Sordello, Matteo, et al.
Published: (2019)
by: Sordello, Matteo, et al.
Published: (2019)
PySCIPOpt-ML: Embedding Trained Machine Learning Models into Mixed-Integer Programs
by: Turner, Mark, et al.
Published: (2023)
by: Turner, Mark, et al.
Published: (2023)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
Importance Sampling Optimization with Laplace Principle
by: Dragomir, Radu-Alexandru, et al.
Published: (2026)
by: Dragomir, Radu-Alexandru, et al.
Published: (2026)
Optimization Learning
by: Van Hentenryck, Pascal
Published: (2025)
by: Van Hentenryck, Pascal
Published: (2025)
Feasible Pairings for Decentralized Integral Controllability of Non-Square Systems
by: Tong, Yuhao, et al.
Published: (2026)
by: Tong, Yuhao, et al.
Published: (2026)
Exploiting Symmetries in Optimal Quantum Circuit Design
by: de Meijer, Frank, et al.
Published: (2024)
by: de Meijer, Frank, et al.
Published: (2024)
Using Laplace Transform To Optimize the Hallucination of Generation Models
by: Kang, Cheng, et al.
Published: (2026)
by: Kang, Cheng, et al.
Published: (2026)
Policy Optimization for PDE Control with a Warm Start
by: Zhang, Xiangyuan, et al.
Published: (2024)
by: Zhang, Xiangyuan, et al.
Published: (2024)
ODE-based Learning to Optimize
by: Xie, Zhonglin, et al.
Published: (2024)
by: Xie, Zhonglin, et al.
Published: (2024)
The Internal Model Principle of Time-Varying Optimization
by: Bianchin, Gianluca, et al.
Published: (2024)
by: Bianchin, Gianluca, et al.
Published: (2024)
Extending Parametric Model Embedding with Physical Information for Design-space Dimensionality Reduction in Shape Optimization
by: Serani, Andrea, et al.
Published: (2025)
by: Serani, Andrea, et al.
Published: (2025)
Embedded State Estimation for Optimization of Cislunar Space Domain Awareness Constellation Design
by: Clareson, Thomas H., et al.
Published: (2024)
by: Clareson, Thomas H., et al.
Published: (2024)
Unifying Controller Design for Stabilizing Nonlinear Systems with Norm-Bounded Control Inputs
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Complexity Bounds for Smooth Multiobjective Optimization
by: Sampaio, Phillipe R.
Published: (2025)
by: Sampaio, Phillipe R.
Published: (2025)
Compact Optimality Verification for Optimization Proxies
by: Chen, Wenbo, et al.
Published: (2024)
by: Chen, Wenbo, et al.
Published: (2024)
Magazine Supply Optimization: a Case-study
by: Nguyen, Duong, et al.
Published: (2024)
by: Nguyen, Duong, et al.
Published: (2024)
Similar Items
-
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
by: Lau, Tim Tsz-Kit, et al.
Published: (2025) -
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
by: Yadav, Sarthak, et al.
Published: (2025) -
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024) -
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
by: Du, Zhehang, et al.
Published: (2026) -
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
by: Su, Weijie
Published: (2025)