LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Rui, Liu, Xiang, Diao, Shizhe, Pi, Renjie, Zhang, Jipeng, Han, Chi, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Effective Bilevel Optimization via Minimax Reformulation
by: Wang, Xiaoyu, et al.
Published: (2023)
by: Wang, Xiaoyu, et al.
Published: (2023)
ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
by: Wang, Ruida, et al.
Published: (2024)
by: Wang, Ruida, et al.
Published: (2024)
Hyperparameter Optimization for Large Language Model Instruction-Tuning
by: Tribes, Christophe, et al.
Published: (2023)
by: Tribes, Christophe, et al.
Published: (2023)
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
by: Han, Tianyang, et al.
Published: (2024)
by: Han, Tianyang, et al.
Published: (2024)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
by: Gautam, Tanmay, et al.
Published: (2024)
by: Gautam, Tanmay, et al.
Published: (2024)
MA-LoT: Model-Collaboration Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving
by: Wang, Ruida, et al.
Published: (2025)
by: Wang, Ruida, et al.
Published: (2025)
COS-DPO: Conditioned One-Shot Multi-Objective Fine-Tuning Framework
by: Ren, Yinuo, et al.
Published: (2024)
by: Ren, Yinuo, et al.
Published: (2024)
TAGCOS: Task-agnostic Gradient Clustered Coreset Selection for Instruction Tuning Data
by: Zhang, Jipeng, et al.
Published: (2024)
by: Zhang, Jipeng, et al.
Published: (2024)
Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
Parameter-Efficient Subspace Optimization for LLM Fine-Tuning
by: Lou, Yuchen, et al.
Published: (2025)
by: Lou, Yuchen, et al.
Published: (2025)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
by: Fernando, Heshan, et al.
Published: (2024)
by: Fernando, Heshan, et al.
Published: (2024)
Importance Sampling Optimization with Laplace Principle
by: Dragomir, Radu-Alexandru, et al.
Published: (2026)
by: Dragomir, Radu-Alexandru, et al.
Published: (2026)
Sample-Efficient Counterfactual Tuning for Compressor Pressure Control
by: Guerrero, Margarita A., et al.
Published: (2025)
by: Guerrero, Margarita A., et al.
Published: (2025)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM Training
by: Refael, Yehonathan, et al.
Published: (2025)
by: Refael, Yehonathan, et al.
Published: (2025)
Active Prompting with Chain-of-Thought for Large Language Models
by: Diao, Shizhe, et al.
Published: (2023)
by: Diao, Shizhe, et al.
Published: (2023)
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024)
by: Liu, Yuxing, et al.
Published: (2024)
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
by: Parfenov, Valery, et al.
Published: (2026)
by: Parfenov, Valery, et al.
Published: (2026)
Sampling Observability for Heat Equations with Memory
by: Ma, Lingying, et al.
Published: (2024)
by: Ma, Lingying, et al.
Published: (2024)
MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
by: Diao, Shizhe, et al.
Published: (2023)
by: Diao, Shizhe, et al.
Published: (2023)
Solving General Natural-Language-Description Optimization Problems with Large Language Models
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
Work Smarter...Not Harder: Efficient Minimization of Dependency Length in SOV Languages
by: Ranjan, Sidharth, et al.
Published: (2024)
by: Ranjan, Sidharth, et al.
Published: (2024)
Importance Sampling in Expensive Finite-Sum Optimization via Contextual Bandit Methods
by: Menickelly, Matt
Published: (2026)
by: Menickelly, Matt
Published: (2026)
Layerwise goal-oriented adaptivity for neural ODEs: an optimal control perspective
by: Hintermüller, Michael, et al.
Published: (2026)
by: Hintermüller, Michael, et al.
Published: (2026)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
DPZero: Private Fine-Tuning of Language Models without Backpropagation
by: Zhang, Liang, et al.
Published: (2023)
by: Zhang, Liang, et al.
Published: (2023)
Leveraging Large Language Models for Solving Rare MIP Challenges
by: Wang, Teng, et al.
Published: (2024)
by: Wang, Teng, et al.
Published: (2024)
Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient Noise
by: Pan, Rui, et al.
Published: (2023)
by: Pan, Rui, et al.
Published: (2023)
Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning
by: Qiu, Shuang, et al.
Published: (2024)
by: Qiu, Shuang, et al.
Published: (2024)
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
Efficient Tail-Aware Generative Optimization via Flow Model Fine-Tuning
by: Wang, Zifan, et al.
Published: (2026)
by: Wang, Zifan, et al.
Published: (2026)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
An Adaptive and Stability-Promoting Layerwise Training Approach for Sparse Deep Neural Network Architecture
by: Krishnanunni, C G, et al.
Published: (2022)
by: Krishnanunni, C G, et al.
Published: (2022)
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
by: Liu, Yuxing, et al.
Published: (2025)
by: Liu, Yuxing, et al.
Published: (2025)
A Reliability Theory of Compromise Decisions for Large-Scale Stochastic Programs
by: Diao, Shuotao, et al.
Published: (2024)
by: Diao, Shuotao, et al.
Published: (2024)
Similar Items
-
Effective Bilevel Optimization via Minimax Reformulation
by: Wang, Xiaoyu, et al.
Published: (2023) -
ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting
by: Pan, Rui, et al.
Published: (2024) -
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
by: Pan, Rui, et al.
Published: (2025) -
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
by: Wang, Ruida, et al.
Published: (2024) -
Hyperparameter Optimization for Large Language Model Instruction-Tuning
by: Tribes, Christophe, et al.
Published: (2023)