Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Jia-Hong, Shen, Yixian, Zhu, Hongyi, Rudinac, Stevan, Kanoulas, Evangelos |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
di: Huang, Jia-Hong, et al.
Pubblicazione: (2026)
di: Huang, Jia-Hong, et al.
Pubblicazione: (2026)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024)
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024)
MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection
di: Shen, Yixian, et al.
Pubblicazione: (2024)
di: Shen, Yixian, et al.
Pubblicazione: (2024)
MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection
di: Shen, Yixian, et al.
Pubblicazione: (2025)
di: Shen, Yixian, et al.
Pubblicazione: (2025)
Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models
di: Zhu, Hongyi, et al.
Pubblicazione: (2024)
di: Zhu, Hongyi, et al.
Pubblicazione: (2024)
A Novel Evaluation Framework for Image2Text Generation
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024)
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024)
VL-KGE: Vision-Language Models Meet Knowledge Graph Embeddings
di: Efthymiou, Athanasios, et al.
Pubblicazione: (2026)
di: Efthymiou, Athanasios, et al.
Pubblicazione: (2026)
Lotus: Efficient LLM Training by Randomized Low-Rank Gradient Projection with Adaptive Subspace Switching
di: Miao, Tianhao, et al.
Pubblicazione: (2026)
di: Miao, Tianhao, et al.
Pubblicazione: (2026)
GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection
di: Su, DiJia, et al.
Pubblicazione: (2025)
di: Su, DiJia, et al.
Pubblicazione: (2025)
A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding
di: Wang, Shuai, et al.
Pubblicazione: (2026)
di: Wang, Shuai, et al.
Pubblicazione: (2026)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
di: Xu, Conglong, et al.
Pubblicazione: (2025)
di: Xu, Conglong, et al.
Pubblicazione: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2024)
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2024)
Optimizing Numerical Estimation and Operational Efficiency in the Legal Domain through Large Language Models
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024)
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024)
Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?
di: Chen, Xi, et al.
Pubblicazione: (2024)
di: Chen, Xi, et al.
Pubblicazione: (2024)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
di: Yi, Qingao, et al.
Pubblicazione: (2025)
di: Yi, Qingao, et al.
Pubblicazione: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
di: Yu, Xiaoming, et al.
Pubblicazione: (2026)
di: Yu, Xiaoming, et al.
Pubblicazione: (2026)
COAP: Memory-Efficient Training with Correlation-Aware Gradient Projection
di: Xiao, Jinqi, et al.
Pubblicazione: (2024)
di: Xiao, Jinqi, et al.
Pubblicazione: (2024)
Unbiased Gradient Low-Rank Projection
di: Pan, Rui, et al.
Pubblicazione: (2025)
di: Pan, Rui, et al.
Pubblicazione: (2025)
Gradient Multi-Normalization for Stateless and Scalable LLM Training
di: Scetbon, Meyer, et al.
Pubblicazione: (2025)
di: Scetbon, Meyer, et al.
Pubblicazione: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
di: Yan, Xianglong, et al.
Pubblicazione: (2026)
di: Yan, Xianglong, et al.
Pubblicazione: (2026)
FwdLLM: Efficient FedLLM using Forward Gradient
di: Xu, Mengwei, et al.
Pubblicazione: (2023)
di: Xu, Mengwei, et al.
Pubblicazione: (2023)
NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning
di: Zhang, Zhi, et al.
Pubblicazione: (2025)
di: Zhang, Zhi, et al.
Pubblicazione: (2025)
PrunedLoRA: Robust Gradient-Based structured pruning for Low-rank Adaptation in Fine-tuning
di: Yu, Xin, et al.
Pubblicazione: (2025)
di: Yu, Xin, et al.
Pubblicazione: (2025)
Dynamic Low-rank Approximation of Full-Matrix Preconditioner for Training Generalized Linear Models
di: Matveeva, Tatyana, et al.
Pubblicazione: (2025)
di: Matveeva, Tatyana, et al.
Pubblicazione: (2025)
When Losses Align: Gradient-Based Composite Loss Weighting for Efficient Pretraining
di: Karpukhin, Ivan, et al.
Pubblicazione: (2026)
di: Karpukhin, Ivan, et al.
Pubblicazione: (2026)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
di: Dong, Harry, et al.
Pubblicazione: (2025)
di: Dong, Harry, et al.
Pubblicazione: (2025)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
di: Lv, Xingtai, et al.
Pubblicazione: (2024)
di: Lv, Xingtai, et al.
Pubblicazione: (2024)
Beyond Low-rank Decomposition: A Shortcut Approach for Efficient On-Device Learning
di: Nguyen, Le-Trung, et al.
Pubblicazione: (2025)
di: Nguyen, Le-Trung, et al.
Pubblicazione: (2025)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
di: Ye, Hao, et al.
Pubblicazione: (2026)
di: Ye, Hao, et al.
Pubblicazione: (2026)
Approximated Orthogonal Projection Unit: Stabilizing Regression Network Training Using Natural Gradient
di: Wang, Shaoqi, et al.
Pubblicazione: (2024)
di: Wang, Shaoqi, et al.
Pubblicazione: (2024)
4-bit Shampoo for Memory-Efficient Network Training
di: Wang, Sike, et al.
Pubblicazione: (2024)
di: Wang, Sike, et al.
Pubblicazione: (2024)
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
di: Lu, Yishun, et al.
Pubblicazione: (2025)
di: Lu, Yishun, et al.
Pubblicazione: (2025)
Tensor Train Low-rank Approximation (TT-LoRA): Democratizing AI with Accelerated LLMs
di: Anjum, Afia, et al.
Pubblicazione: (2024)
di: Anjum, Afia, et al.
Pubblicazione: (2024)
Training Data Selection with Gradient Orthogonality for Efficient Domain Adaptation
di: Zhang, Xiyang, et al.
Pubblicazione: (2026)
di: Zhang, Xiyang, et al.
Pubblicazione: (2026)
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
di: Wu, Bo, et al.
Pubblicazione: (2025)
di: Wu, Bo, et al.
Pubblicazione: (2025)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
di: Chen, Zhipeng, et al.
Pubblicazione: (2026)
di: Chen, Zhipeng, et al.
Pubblicazione: (2026)
Enhancing DP-SGD through Non-monotonous Adaptive Scaling Gradient Weight
di: Huang, Tao, et al.
Pubblicazione: (2024)
di: Huang, Tao, et al.
Pubblicazione: (2024)
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning
di: Zhang, Yixian, et al.
Pubblicazione: (2025)
di: Zhang, Yixian, et al.
Pubblicazione: (2025)
DeltaLLM: Compress LLMs with Low-Rank Deltas between Shared Weights
di: Mikaelyan, Liana, et al.
Pubblicazione: (2025)
di: Mikaelyan, Liana, et al.
Pubblicazione: (2025)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
di: Melo, Luckeciano C., et al.
Pubblicazione: (2025)
di: Melo, Luckeciano C., et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
di: Huang, Jia-Hong, et al.
Pubblicazione: (2026) -
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024) -
MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection
di: Shen, Yixian, et al.
Pubblicazione: (2024) -
MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection
di: Shen, Yixian, et al.
Pubblicazione: (2025) -
Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models
di: Zhu, Hongyi, et al.
Pubblicazione: (2024)