LoRA-Switch: Boosting the Efficiency of Dynamic LLM Adapters via System-Algorithm Co-design
Fuente:
arXiv
Saved in:
| Main Authors: | Kong, Rui, Li, Qiyang, Fang, Xinyu, Feng, Qingtian, He, Qingfeng, Dong, Yazhu, Wang, Weijun, Li, Yuanchun, Kong, Linghe, Liu, Yunxin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Empower Vision Applications with LoRA LMM
by: Mi, Liang, et al.
Published: (2024)
by: Mi, Liang, et al.
Published: (2024)
SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
by: Kong, Rui, et al.
Published: (2023)
by: Kong, Rui, et al.
Published: (2023)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
by: Sheng, Ying, et al.
Published: (2023)
by: Sheng, Ying, et al.
Published: (2023)
AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization
by: Li, Qiyang, et al.
Published: (2026)
by: Li, Qiyang, et al.
Published: (2026)
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs
by: Ye, Zhengmao, et al.
Published: (2023)
by: Ye, Zhengmao, et al.
Published: (2023)
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving
by: Li, Shaoang, et al.
Published: (2026)
by: Li, Shaoang, et al.
Published: (2026)
AuthenLoRA: Entangling Stylization with Imperceptible Watermarks for Copyright-Secure LoRA Adapters
by: Shi, Fangming, et al.
Published: (2025)
by: Shi, Fangming, et al.
Published: (2025)
LoRACode: LoRA Adapters for Code Embeddings
by: Chaturvedi, Saumya, et al.
Published: (2025)
by: Chaturvedi, Saumya, et al.
Published: (2025)
Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
by: Shen, Yixin
Published: (2025)
by: Shen, Yixin
Published: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
DragLoRA: Online Optimization of LoRA Adapters for Drag-based Image Editing in Diffusion Model
by: Xia, Siwei, et al.
Published: (2025)
by: Xia, Siwei, et al.
Published: (2025)
Weight space Detection of Backdoors in LoRA Adapters
by: Merenciano, David Puertolas, et al.
Published: (2026)
by: Merenciano, David Puertolas, et al.
Published: (2026)
LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference
by: Rajabzadeh, Hossein, et al.
Published: (2026)
by: Rajabzadeh, Hossein, et al.
Published: (2026)
LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
Effective LoRA Adapter Routing using Task Representations
by: Dhasade, Akash, et al.
Published: (2026)
by: Dhasade, Akash, et al.
Published: (2026)
SLAD : Shared LoRA Adapters for Task Specific Distillation
by: Bensaid, Reda, et al.
Published: (2026)
by: Bensaid, Reda, et al.
Published: (2026)
Improving SAM for Camouflaged Object Detection via Dual Stream Adapters
by: Liu, Jiaming, et al.
Published: (2025)
by: Liu, Jiaming, et al.
Published: (2025)
Cross-LoRA: A Data-Free LoRA Transfer Framework across Heterogeneous LLMs
by: Xia, Feifan, et al.
Published: (2025)
by: Xia, Feifan, et al.
Published: (2025)
PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRA
by: Wang, Sheng, et al.
Published: (2024)
by: Wang, Sheng, et al.
Published: (2024)
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA Serving
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
by: Li, Allison, et al.
Published: (2025)
by: Li, Allison, et al.
Published: (2025)
SHE-LoRA: Selective Homomorphic Encryption for Federated Tuning with Heterogeneous LoRA
by: Liu, Jianmin, et al.
Published: (2025)
by: Liu, Jianmin, et al.
Published: (2025)
LoRA-PAR: A Flexible Dual-System LoRA Partitioning Approach to Efficient LLM Fine-Tuning
by: Huang, Yining, et al.
Published: (2025)
by: Huang, Yining, et al.
Published: (2025)
BoostLoRA: Growing Effective Rank by Boosting Adapters
by: Anantha, Raviteja, et al.
Published: (2026)
by: Anantha, Raviteja, et al.
Published: (2026)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
Vision as LoRA
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters
by: Ahn, Beomjin, et al.
Published: (2026)
by: Ahn, Beomjin, et al.
Published: (2026)
HypeLoRA: Hyper-Network-Generated LoRA Adapters for Calibrated Language Model Fine-Tuning
by: Trojan, Bartosz, et al.
Published: (2026)
by: Trojan, Bartosz, et al.
Published: (2026)
LoRA Meets Dropout under a Unified Framework
by: Wang, Sheng, et al.
Published: (2024)
by: Wang, Sheng, et al.
Published: (2024)
Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment
by: Fan, Chenghao, et al.
Published: (2025)
by: Fan, Chenghao, et al.
Published: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
by: Chen, Hongyu, et al.
Published: (2026)
by: Chen, Hongyu, et al.
Published: (2026)
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
by: Pletenev, Sergey, et al.
Published: (2025)
by: Pletenev, Sergey, et al.
Published: (2025)
Boosting Robust AIGI Detection with LoRA-based Pairwise Training
by: Xia, Ruiyang, et al.
Published: (2026)
by: Xia, Ruiyang, et al.
Published: (2026)
Improving Fisher Information Estimation and Efficiency for LoRA-based LLM Unlearning
by: Kim, Yejin, et al.
Published: (2025)
by: Kim, Yejin, et al.
Published: (2025)
Parametric Retrieval-Augmented Generation using Latent Routing of LoRA Adapters
by: Su, Zhan, et al.
Published: (2025)
by: Su, Zhan, et al.
Published: (2025)
Run LoRA Run: Faster and Lighter LoRA Implementations
by: Cherniuk, Daria, et al.
Published: (2023)
by: Cherniuk, Daria, et al.
Published: (2023)
FREE-Switch: Frequency-based Dynamic LoRA Switch for Style Transfer
by: Zheng, Shenghe, et al.
Published: (2026)
by: Zheng, Shenghe, et al.
Published: (2026)
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
by: Li, Wenbing, et al.
Published: (2025)
by: Li, Wenbing, et al.
Published: (2025)
Similar Items
-
Empower Vision Applications with LoRA LMM
by: Mi, Liang, et al.
Published: (2024) -
SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
by: Kong, Rui, et al.
Published: (2023) -
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
by: Sheng, Ying, et al.
Published: (2023) -
AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization
by: Li, Qiyang, et al.
Published: (2026) -
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs
by: Ye, Zhengmao, et al.
Published: (2023)