GradientStabilizer:Fix the Norm, Not the Gradient
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Tianjin, Wang, Zhangyang, Hu, Haotian, Zhang, Zhenyu, Jin, Gaojie, Li, Xiang, Shen, Li, Shang, Jiaxing, Chen, Tianlong, Li, Ke, Liu, Lu, Wen, Qingsong, Liu, Shiwei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
por: Huang, Tianjin, et al.
Publicado: (2025)
por: Huang, Tianjin, et al.
Publicado: (2025)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
por: Zhang, Zhenyu, et al.
Publicado: (2024)
por: Zhang, Zhenyu, et al.
Publicado: (2024)
OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents
por: Li, Xinyu, et al.
Publicado: (2026)
por: Li, Xinyu, et al.
Publicado: (2026)
(PASS) Visual Prompt Locates Good Structure Sparsity through a Recurrent HyperNetwork
por: Huang, Tianjin, et al.
Publicado: (2024)
por: Huang, Tianjin, et al.
Publicado: (2024)
POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
por: Li, Xinyu, et al.
Publicado: (2025)
por: Li, Xinyu, et al.
Publicado: (2025)
Visual Prompting Upgrades Neural Network Sparsification: A Data-Model Perspective
por: Jin, Can, et al.
Publicado: (2023)
por: Jin, Can, et al.
Publicado: (2023)
Principal Eigenvalue Regularization for Improved Worst-Class Certified Robustness of Smoothed Classifiers
por: Jin, Gaojie, et al.
Publicado: (2025)
por: Jin, Gaojie, et al.
Publicado: (2025)
Enhancing Adversarial Training via Reweighting Optimization Trajectory
por: Huang, Tianjin, et al.
Publicado: (2023)
por: Huang, Tianjin, et al.
Publicado: (2023)
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
por: Jaiswal, Ajay, et al.
Publicado: (2024)
por: Jaiswal, Ajay, et al.
Publicado: (2024)
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
por: Jin, Gaojie, et al.
Publicado: (2026)
por: Jin, Gaojie, et al.
Publicado: (2026)
Enhancing Robust Fairness via Confusional Spectral Regularization
por: Jin, Gaojie, et al.
Publicado: (2025)
por: Jin, Gaojie, et al.
Publicado: (2025)
Dynamic Proximal Gradient Algorithms for Schatten-$p$ Quasi-Norm Regularized Problems
por: Shen, Weiping, et al.
Publicado: (2026)
por: Shen, Weiping, et al.
Publicado: (2026)
Dual-Kernel Adapter: Expanding Spatial Horizons for Data-Constrained Medical Image Analysis
por: Zhu, Ziquan, et al.
Publicado: (2026)
por: Zhu, Ziquan, et al.
Publicado: (2026)
Confusion-Aware Spectral Regularizer for Long-Tailed Recognition
por: Zhu, Ziquan, et al.
Publicado: (2026)
por: Zhu, Ziquan, et al.
Publicado: (2026)
You Can Have Better Graph Neural Networks by Not Training Weights at All: Finding Untrained GNNs Tickets
por: Huang, Tianjin, et al.
Publicado: (2022)
por: Huang, Tianjin, et al.
Publicado: (2022)
Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech Recognition
por: Zhou, Nanjun, et al.
Publicado: (2025)
por: Zhou, Nanjun, et al.
Publicado: (2025)
Gradient Doping Strategy for Sn─Pb Mixed Perovskite Solar Cells with High Efficiency and Stability
por: Haotian Zhang, et al.
Publicado: (2025)
por: Haotian Zhang, et al.
Publicado: (2025)
Q-Newton: Hybrid Quantum-Classical Scheduling for Accelerating Neural Network Training with Newton's Gradient Descent
por: Li, Pingzhi, et al.
Publicado: (2024)
por: Li, Pingzhi, et al.
Publicado: (2024)
Elementary Analysis of Policy Gradient Methods
por: Liu, Jiacai, et al.
Publicado: (2024)
por: Liu, Jiacai, et al.
Publicado: (2024)
Discriminability-Driven Spatial-Channel Selection with Gradient Norm for Drone Signal OOD Detection
por: Feng, Chuhan, et al.
Publicado: (2026)
por: Feng, Chuhan, et al.
Publicado: (2026)
LOST: Low-rank and Sparse Pre-training for Large Language Models
por: Li, Jiaxi, et al.
Publicado: (2025)
por: Li, Jiaxi, et al.
Publicado: (2025)
A Unified Analysis of Stochastic Gradient Descent with Arbitrary Data Permutations and Beyond
por: Li, Yipeng, et al.
Publicado: (2025)
por: Li, Yipeng, et al.
Publicado: (2025)
Towards A Unified PAC-Bayesian Framework for Norm-based Generalization Bounds
por: Yi, Xinping, et al.
Publicado: (2026)
por: Yi, Xinping, et al.
Publicado: (2026)
Automate Knowledge Concept Tagging on Math Questions with LLMs
por: Li, Hang, et al.
Publicado: (2024)
por: Li, Hang, et al.
Publicado: (2024)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever
por: Li, Hang, et al.
Publicado: (2024)
por: Li, Hang, et al.
Publicado: (2024)
Knowledge Tagging with Large Language Model based Multi-Agent System
por: Li, Hang, et al.
Publicado: (2024)
por: Li, Hang, et al.
Publicado: (2024)
Simple Fabrication of Porous Nanocomposites with Gradient Composition for High‐Performance Terahertz Absorbers
por: Junxiao Liu, et al.
Publicado: (2025)
por: Junxiao Liu, et al.
Publicado: (2025)
Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?
por: Li, Pengxiang, et al.
Publicado: (2026)
por: Li, Pengxiang, et al.
Publicado: (2026)
FedBCD:Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning
por: Liu, Junkang, et al.
Publicado: (2026)
por: Liu, Junkang, et al.
Publicado: (2026)
Defense Against Adversarial Attacks on No-Reference Image Quality Models with Gradient Norm Regularization
por: Liu, Yujia, et al.
Publicado: (2024)
por: Liu, Yujia, et al.
Publicado: (2024)
Gradient Catastrophe for Solutions to the Hyperbolic Navier-Stokes Equations
por: Zhao, Qingsong
Publicado: (2026)
por: Zhao, Qingsong
Publicado: (2026)
Gradient Catastrophe for Solutions to the Conservation Laws with Source Term
por: Zhao, Qingsong
Publicado: (2026)
por: Zhao, Qingsong
Publicado: (2026)
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
por: Zhao, Jiawei, et al.
Publicado: (2024)
por: Zhao, Jiawei, et al.
Publicado: (2024)
SARC: Sentiment-Augmented Deep Role Clustering for Fake News Detection
por: Wang, Jingqing, et al.
Publicado: (2025)
por: Wang, Jingqing, et al.
Publicado: (2025)
DP-FedPGN: Finding Global Flat Minima for Differentially Private Federated Learning via Penalizing Gradient Norm
por: Liu, Junkang, et al.
Publicado: (2025)
por: Liu, Junkang, et al.
Publicado: (2025)
Type-II Saddles and Probabilistic Stability of Stochastic Gradient Descent
por: Ziyin, Liu, et al.
Publicado: (2023)
por: Ziyin, Liu, et al.
Publicado: (2023)
$\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space
por: Wang, Peihao, et al.
Publicado: (2026)
por: Wang, Peihao, et al.
Publicado: (2026)
Scaling Textual Gradients via Sampling-Based Momentum
por: Ding, Zixin, et al.
Publicado: (2025)
por: Ding, Zixin, et al.
Publicado: (2025)
Continuous-tone Simple Points: An $\ell_0$-Norm of Cyclic Gradient for Topology-Preserving Data-Driven Image Segmentation
por: Li, Wenxiao, et al.
Publicado: (2026)
por: Li, Wenxiao, et al.
Publicado: (2026)
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
por: Wen, Xuexiang, et al.
Publicado: (2026)
por: Wen, Xuexiang, et al.
Publicado: (2026)
Ejemplares similares
-
SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
por: Huang, Tianjin, et al.
Publicado: (2025) -
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
por: Zhang, Zhenyu, et al.
Publicado: (2024) -
OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents
por: Li, Xinyu, et al.
Publicado: (2026) -
(PASS) Visual Prompt Locates Good Structure Sparsity through a Recurrent HyperNetwork
por: Huang, Tianjin, et al.
Publicado: (2024) -
POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
por: Li, Xinyu, et al.
Publicado: (2025)