Saved in:
| Main Authors: | Liu, Yizhou, Liu, Ziming, Gore, Jeff |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2501.12243 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Superposition Yields Robust Neural Scaling
by: Liu, Yizhou, et al.
Published: (2025)
by: Liu, Yizhou, et al.
Published: (2025)
First-Order Sparse Convex Optimization: Better Rates with Sparse Updates
by: Garber, Dan
Published: (2025)
by: Garber, Dan
Published: (2025)
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback
by: Wang, Yikai, et al.
Published: (2026)
by: Wang, Yikai, et al.
Published: (2026)
Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
by: Hanqing, Liu, et al.
Published: (2026)
by: Hanqing, Liu, et al.
Published: (2026)
Convergence for Discrete Parameter Update Schemes
by: Wilson, Paul, et al.
Published: (2025)
by: Wilson, Paul, et al.
Published: (2025)
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
by: Liu, Hong, et al.
Published: (2023)
by: Liu, Hong, et al.
Published: (2023)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
by: Gautam, Tanmay, et al.
Published: (2024)
by: Gautam, Tanmay, et al.
Published: (2024)
A Unified Understanding of Offline Data Selection and Online Self-refining Generation for Post-training LLMs
by: Xiao, Quan, et al.
Published: (2025)
by: Xiao, Quan, et al.
Published: (2025)
Distributional Surgery for Language Model Activations
by: Nguyen, Bao, et al.
Published: (2025)
by: Nguyen, Bao, et al.
Published: (2025)
SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM Training
by: Refael, Yehonathan, et al.
Published: (2025)
by: Refael, Yehonathan, et al.
Published: (2025)
Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices
by: Zhao, Pengxiang, et al.
Published: (2024)
by: Zhao, Pengxiang, et al.
Published: (2024)
COS-DPO: Conditioned One-Shot Multi-Objective Fine-Tuning Framework
by: Ren, Yinuo, et al.
Published: (2024)
by: Ren, Yinuo, et al.
Published: (2024)
Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
by: Bu, Zhiqi, et al.
Published: (2026)
by: Bu, Zhiqi, et al.
Published: (2026)
Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
by: Kunstner, Frederik, et al.
Published: (2024)
by: Kunstner, Frederik, et al.
Published: (2024)
When and How Unlabeled Data Provably Improve In-Context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
by: Xu, Yizhou, et al.
Published: (2026)
by: Xu, Yizhou, et al.
Published: (2026)
Online Optimization Perspective on First-Order and Zero-Order Decentralized Nonsmooth Nonconvex Stochastic Optimization
by: Sahinoglu, Emre, et al.
Published: (2024)
by: Sahinoglu, Emre, et al.
Published: (2024)
Fully First-Order Algorithms for Online Bilevel Optimization
by: Jia, Tingkai, et al.
Published: (2026)
by: Jia, Tingkai, et al.
Published: (2026)
On the Complexity of First-Order Methods in Stochastic Bilevel Optimization
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
First-Order Methods for Linearly Constrained Bilevel Optimization
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
A Study of Condition Numbers for First-Order Optimization
by: Guille-Escuret, Charles, et al.
Published: (2020)
by: Guille-Escuret, Charles, et al.
Published: (2020)
Accelerated Fully First-Order Methods for Bilevel and Minimax Optimization
by: Li, Chris Junchi
Published: (2024)
by: Li, Chris Junchi
Published: (2024)
Batched First-Order Methods for Parallel LP Solving in MIP
by: Blin, Nicolas, et al.
Published: (2026)
by: Blin, Nicolas, et al.
Published: (2026)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
by: Kharrat, Salma, et al.
Published: (2024)
by: Kharrat, Salma, et al.
Published: (2024)
ControlAgent: Automating Control System Design via Novel Integration of LLM Agents and Domain Expertise
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning
by: Sherman, Uri, et al.
Published: (2025)
by: Sherman, Uri, et al.
Published: (2025)
First Order Methods with Markovian Noise: from Acceleration to Variational Inequalities
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
On Penalty Methods for Nonconvex Bilevel Optimization and First-Order Stochastic Approximation
by: Kwon, Jeongyeol, et al.
Published: (2023)
by: Kwon, Jeongyeol, et al.
Published: (2023)
A New First-Order Meta-Learning Algorithm with Convergence Guarantees
by: Chayti, El Mahdi, et al.
Published: (2024)
by: Chayti, El Mahdi, et al.
Published: (2024)
Causal LLM Routing: End-to-End Regret Minimization from Observational Data
by: Tsiourvas, Asterios, et al.
Published: (2025)
by: Tsiourvas, Asterios, et al.
Published: (2025)
DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
by: Gurses, Selcuk, et al.
Published: (2025)
by: Gurses, Selcuk, et al.
Published: (2025)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
by: Fernando, Heshan, et al.
Published: (2024)
by: Fernando, Heshan, et al.
Published: (2024)
Solving General Natural-Language-Description Optimization Problems with Large Language Models
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
by: Huang, Xinmeng, et al.
Published: (2024)
by: Huang, Xinmeng, et al.
Published: (2024)
Reinforcement Learning from Human Feedback with Active Queries
by: Ji, Kaixuan, et al.
Published: (2024)
by: Ji, Kaixuan, et al.
Published: (2024)
Reward Collapse in Aligning Large Language Models
by: Song, Ziang, et al.
Published: (2023)
by: Song, Ziang, et al.
Published: (2023)
Similar Items
-
Superposition Yields Robust Neural Scaling
by: Liu, Yizhou, et al.
Published: (2025) -
First-Order Sparse Convex Optimization: Better Rates with Sparse Updates
by: Garber, Dan
Published: (2025) -
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback
by: Wang, Yikai, et al.
Published: (2026) -
Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
by: Hanqing, Liu, et al.
Published: (2026) -
Convergence for Discrete Parameter Update Schemes
by: Wilson, Paul, et al.
Published: (2025)