WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle Points
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Dongyue, Liu, Zechun, Yi, Kai, Zhang, Zhenshuo, Zhao, Changsheng, Krishnamoorthi, Raghuraman, Khaitan, Harshit, Zhang, Hongyang R., Li, Steven |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Iterative Proximal-Minimization for Computing Saddle Points with Fixed Index
by: Gu, Shuting, et al.
Published: (2025)
by: Gu, Shuting, et al.
Published: (2025)
Accelerated Primal-Dual Proximal Gradient Splitting Methods for Convex-Concave Saddle-Point Problems
by: Luo, Hao
Published: (2024)
by: Luo, Hao
Published: (2024)
A Framework for the Solution of Tree-Coupled Saddle-Point Systems
by: Hansknecht, Christoph, et al.
Published: (2024)
by: Hansknecht, Christoph, et al.
Published: (2024)
A Derivative-Free Saddle-search Algorithm With Linear Convergence Rate
by: Du, Qiang, et al.
Published: (2026)
by: Du, Qiang, et al.
Published: (2026)
A Generalized Primal-Dual Correction Method for Saddle-Point Problems with a Nonlinear Coupling Operator
by: Wang, Sai, et al.
Published: (2023)
by: Wang, Sai, et al.
Published: (2023)
Consensus-Based Optimization for Saddle Point Problems
by: Huang, Hui, et al.
Published: (2022)
by: Huang, Hui, et al.
Published: (2022)
JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates
by: Yi, Kai, et al.
Published: (2026)
by: Yi, Kai, et al.
Published: (2026)
Optimal Design of Broadband Absorbers with Multiple Plasmonic Nanoparticles via Reduced Basis Method
by: Gao, Yu, et al.
Published: (2025)
by: Gao, Yu, et al.
Published: (2025)
On the convergence of iterative regularization method assisted by the graph Laplacian with early stopping
by: Bajpai, Harshit, et al.
Published: (2025)
by: Bajpai, Harshit, et al.
Published: (2025)
Robust Accelerated Primal-Dual Methods for Computing Saddle Points
by: Zhang, Xuan, et al.
Published: (2021)
by: Zhang, Xuan, et al.
Published: (2021)
Quantization Avoids Saddle Points in Distributed Optimization
by: Bo, Yanan, et al.
Published: (2024)
by: Bo, Yanan, et al.
Published: (2024)
Accelerating operator Sinkhorn iteration with overrelaxation
by: Soma, Tasuku, et al.
Published: (2024)
by: Soma, Tasuku, et al.
Published: (2024)
Generalized Composed Alternating Relaxed Projection Algorithm for Two-Set Feasibility Problem
by: Li, Xinxin, et al.
Published: (2026)
by: Li, Xinxin, et al.
Published: (2026)
One-Sided Matrix Completion from Ultra-Sparse Samples
by: Zhang, Hongyang R., et al.
Published: (2026)
by: Zhang, Hongyang R., et al.
Published: (2026)
A Lyapunov Analysis of Accelerated PDHG Algorithms
by: Zeng, Xueying, et al.
Published: (2024)
by: Zeng, Xueying, et al.
Published: (2024)
Direct Spectral Acceleration of First-Order Methods for Saddle Point Problems with Bilinear Coupling
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
Scalable Acceleration for Classification-Based Derivative-Free Optimization
by: Han, Tianyi, et al.
Published: (2023)
by: Han, Tianyi, et al.
Published: (2023)
MAGPIE: Multilevel-Adaptive-Guided Solver for Ptychographic Phase Retrieval
by: Zhang, Borong, et al.
Published: (2025)
by: Zhang, Borong, et al.
Published: (2025)
Primal-dual Accelerated Mirror-Descent Method for Constrained Bilinear Saddle-Point Problems
by: Li, Weijian, et al.
Published: (2024)
by: Li, Weijian, et al.
Published: (2024)
State-dependent temperature control in Langevin diffusions using numerical exploratory Hamiltonian-Jacobi-Bellman equations
by: Wang, Taorui, et al.
Published: (2026)
by: Wang, Taorui, et al.
Published: (2026)
Extension of Switch Point Algorithm to Boundary-Value Problems
by: Hager, William W.
Published: (2023)
by: Hager, William W.
Published: (2023)
General Procedure to Provide High-Probability Guarantees for Stochastic Saddle Point Problems
by: Li, Dongyang, et al.
Published: (2024)
by: Li, Dongyang, et al.
Published: (2024)
Anderson Acceleration in Nonsmooth Problems: Local Convergence via Active Manifold Identification
by: Li, Kexin, et al.
Published: (2024)
by: Li, Kexin, et al.
Published: (2024)
The Dynamical Anatomy of Anderson Acceleration:From Adaptive Momentum to Variable-Mass ODEs
by: Chen, Kewang, et al.
Published: (2025)
by: Chen, Kewang, et al.
Published: (2025)
Fast and Provable Tensor-Train Format Tensor Completion via Precondtioned Riemannian Gradient Descent
by: Bian, Fengmiao, et al.
Published: (2025)
by: Bian, Fengmiao, et al.
Published: (2025)
Block Acceleration Without Momentum: On Optimal Stepsizes of Block Gradient Descent for Least-Squares
by: Peng, Liangzu, et al.
Published: (2024)
by: Peng, Liangzu, et al.
Published: (2024)
Acceleration Methods
by: d'Aspremont, Alexandre, et al.
Published: (2021)
by: d'Aspremont, Alexandre, et al.
Published: (2021)
TRAFS: A Nonsmooth Convex Optimization Algorithm with $\mathcal{O}\left(\frac{1}ε\right)$ Iteration Complexity
by: Jia, Kai, et al.
Published: (2023)
by: Jia, Kai, et al.
Published: (2023)
A Scalable Interior-Point Gauss-Newton Method for PDE-Constrained Optimization with Bound Constraints
by: Hartland, Tucker, et al.
Published: (2024)
by: Hartland, Tucker, et al.
Published: (2024)
Noise Stability Optimization for Finding Flat Minima: A Hessian-based Regularization Approach
by: Zhang, Hongyang R., et al.
Published: (2023)
by: Zhang, Hongyang R., et al.
Published: (2023)
Nonlinear preconditioned primal-dual method for a class of structured minimax problems
by: Zhang, Lu, et al.
Published: (2024)
by: Zhang, Lu, et al.
Published: (2024)
PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network Training
by: Yang, Shenghao, et al.
Published: (2026)
by: Yang, Shenghao, et al.
Published: (2026)
Symmetry & Critical Points
by: Arjevani, Yossi
Published: (2024)
by: Arjevani, Yossi
Published: (2024)
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
A Natural Primal-Dual Hybrid Gradient Method for Adversarial Neural Network Training on Solving Partial Differential Equations
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
Solutions for Underdetermined Generalized Absolute Value Equations
by: Chen, Cairong, et al.
Published: (2024)
by: Chen, Cairong, et al.
Published: (2024)
A Stochastic Algorithm for Searching Saddle Points with Convergence Guarantee
by: Shi, Baoming, et al.
Published: (2025)
by: Shi, Baoming, et al.
Published: (2025)
Smoothed Moreau-Yosida Tensor Train Approximation of State-constrained Optimization Problems under Uncertainty
by: Antil, Harbir, et al.
Published: (2023)
by: Antil, Harbir, et al.
Published: (2023)
Two-scale neural networks for optimal control of linear convection-dominated equations
by: Liu, Sijing, et al.
Published: (2026)
by: Liu, Sijing, et al.
Published: (2026)
An inertial minimal-deformation-rate framework for shape optimization
by: Chen, Falai, et al.
Published: (2026)
by: Chen, Falai, et al.
Published: (2026)
Similar Items
-
Iterative Proximal-Minimization for Computing Saddle Points with Fixed Index
by: Gu, Shuting, et al.
Published: (2025) -
Accelerated Primal-Dual Proximal Gradient Splitting Methods for Convex-Concave Saddle-Point Problems
by: Luo, Hao
Published: (2024) -
A Framework for the Solution of Tree-Coupled Saddle-Point Systems
by: Hansknecht, Christoph, et al.
Published: (2024) -
A Derivative-Free Saddle-search Algorithm With Linear Convergence Rate
by: Du, Qiang, et al.
Published: (2026) -
A Generalized Primal-Dual Correction Method for Saddle-Point Problems with a Nonlinear Coupling Operator
by: Wang, Sai, et al.
Published: (2023)