A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Xuan, Li, Jichu, Zou, Difan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient
by: Li, Jichu, et al.
Published: (2026)
by: Li, Jichu, et al.
Published: (2026)
AMS-QUANT: Adaptive Mantissa Sharing for Floating-point Quantization
by: Lv, Mengtao, et al.
Published: (2025)
by: Lv, Mengtao, et al.
Published: (2025)
Learning under Quantization for High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2025)
by: Zhang, Dechen, et al.
Published: (2025)
Structured Role-Aware Policy Optimization for Multimodal Reasoning
by: Jiang, Bingqing, et al.
Published: (2026)
by: Jiang, Bingqing, et al.
Published: (2026)
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
by: Tang, Xuan, et al.
Published: (2025)
by: Tang, Xuan, et al.
Published: (2025)
Scaling Laws for Precision in High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2026)
by: Zhang, Dechen, et al.
Published: (2026)
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
On the Memorization of Consistency Distillation for Diffusion Models
by: Jiang, Bingqing, et al.
Published: (2026)
by: Jiang, Bingqing, et al.
Published: (2026)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
by: Xie, Chengxing, et al.
Published: (2024)
by: Xie, Chengxing, et al.
Published: (2024)
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning
by: Zhang, Dechen, et al.
Published: (2025)
by: Zhang, Dechen, et al.
Published: (2025)
Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Expressive Power of ReLU and Step Networks under Floating-Point Operations
by: Park, Yeachan, et al.
Published: (2024)
by: Park, Yeachan, et al.
Published: (2024)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
On the Robustness of Transformers against Context Hijacking for Linear Classification
by: Li, Tianle, et al.
Published: (2025)
by: Li, Tianle, et al.
Published: (2025)
ProFlow: Zero-Shot Physics-Consistent Sampling via Proximal Flow Guidance
by: Yu, Zichao, et al.
Published: (2026)
by: Yu, Zichao, et al.
Published: (2026)
STGAN: Spatial-temporal Graph Autoregression Network for Pavement Distress Deterioration Prediction
by: Tong, Shilin, et al.
Published: (2025)
by: Tong, Shilin, et al.
Published: (2025)
On the Convergence of Continual Learning with Adaptive Methods
by: Han, Seungyub, et al.
Published: (2024)
by: Han, Seungyub, et al.
Published: (2024)
Understanding the Dynamics of Demonstration Conflict in In-Context Learning
by: Jiao, Difan, et al.
Published: (2026)
by: Jiao, Difan, et al.
Published: (2026)
AIS: Adaptive Importance Sampling for Quantized RL
by: Zhou, Jiajun, et al.
Published: (2026)
by: Zhou, Jiajun, et al.
Published: (2026)
Representation Collapsing Problems in Vector Quantization
by: Zhao, Wenhao, et al.
Published: (2024)
by: Zhao, Wenhao, et al.
Published: (2024)
Continuous-Time Analysis of Adaptive Optimization and Normalization
by: Gould, Rhys, et al.
Published: (2024)
by: Gould, Rhys, et al.
Published: (2024)
A Methodology Establishing Linear Convergence of Adaptive Gradient Methods under PL Inequality
by: Chakrabarti, Kushal, et al.
Published: (2024)
by: Chakrabarti, Kushal, et al.
Published: (2024)
FedLion: Faster Adaptive Federated Optimization with Fewer Communication
by: Tang, Zhiwei, et al.
Published: (2024)
by: Tang, Zhiwei, et al.
Published: (2024)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
by: Bouzouad, Meriem, et al.
Published: (2026)
by: Bouzouad, Meriem, et al.
Published: (2026)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
by: Gernigon, Cédric, et al.
Published: (2024)
by: Gernigon, Cédric, et al.
Published: (2024)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
by: Ye, Rongguang, et al.
Published: (2025)
by: Ye, Rongguang, et al.
Published: (2025)
FedEGG: Federated Learning with Explicit Global Guidance
by: Zhai, Kun, et al.
Published: (2024)
by: Zhai, Kun, et al.
Published: (2024)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
by: Zhao, Zhixiong, et al.
Published: (2025)
by: Zhao, Zhixiong, et al.
Published: (2025)
On the Convergence of Experience Replay in Policy Optimization: Characterizing Bias, Variance, and Finite-Time Convergence
by: Zheng, Hua, et al.
Published: (2021)
by: Zheng, Hua, et al.
Published: (2021)
Enhancing Machine Learning Model Efficiency through Quantization and Bit Depth Optimization: A Performance Analysis on Healthcare Data
by: Goswami, Mitul, et al.
Published: (2025)
by: Goswami, Mitul, et al.
Published: (2025)
Auto Researching, not hyperparameter tuning: Convergence Analysis of 10,000 Experiments
by: Li, Xiaoyi
Published: (2026)
by: Li, Xiaoyi
Published: (2026)
Revealing the Attention Floating Mechanism in Masked Diffusion Models
by: Dai, Xin, et al.
Published: (2026)
by: Dai, Xin, et al.
Published: (2026)
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
by: Zhang, Jinhao, et al.
Published: (2025)
by: Zhang, Jinhao, et al.
Published: (2025)
HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization
by: Zagitov, Artur, et al.
Published: (2026)
by: Zagitov, Artur, et al.
Published: (2026)
VP-VAE: Rethinking Vector Quantization via Adaptive Vector Perturbation
by: Zhai, Linwei, et al.
Published: (2026)
by: Zhai, Linwei, et al.
Published: (2026)
Interpretable Operator Learning for Inverse Problems via Adaptive Spectral Filtering: Convergence and Discretization Invariance
by: Dong, Hang-Cheng, et al.
Published: (2026)
by: Dong, Hang-Cheng, et al.
Published: (2026)
Almost Linear Convergence under Minimal Score Assumptions: Quantized Transition Diffusion
by: Huang, Xunpeng, et al.
Published: (2025)
by: Huang, Xunpeng, et al.
Published: (2025)
Similar Items
-
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient
by: Li, Jichu, et al.
Published: (2026) -
AMS-QUANT: Adaptive Mantissa Sharing for Floating-point Quantization
by: Lv, Mengtao, et al.
Published: (2025) -
Learning under Quantization for High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2025) -
Structured Role-Aware Policy Optimization for Multimodal Reasoning
by: Jiang, Bingqing, et al.
Published: (2026) -
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
by: Tang, Xuan, et al.
Published: (2025)