LoRA-Switch: Boosting the Efficiency of Dynamic LLM Adapters via System-Algorithm Co-design
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kong, Rui, Li, Qiyang, Fang, Xinyu, Feng, Qingtian, He, Qingfeng, Dong, Yazhu, Wang, Weijun, Li, Yuanchun, Kong, Linghe, Liu, Yunxin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Empower Vision Applications with LoRA LMM
von: Mi, Liang, et al.
Veröffentlicht: (2024)
von: Mi, Liang, et al.
Veröffentlicht: (2024)
SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
von: Kong, Rui, et al.
Veröffentlicht: (2023)
von: Kong, Rui, et al.
Veröffentlicht: (2023)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization
von: Li, Qiyang, et al.
Veröffentlicht: (2026)
von: Li, Qiyang, et al.
Veröffentlicht: (2026)
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs
von: Ye, Zhengmao, et al.
Veröffentlicht: (2023)
von: Ye, Zhengmao, et al.
Veröffentlicht: (2023)
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving
von: Li, Shaoang, et al.
Veröffentlicht: (2026)
von: Li, Shaoang, et al.
Veröffentlicht: (2026)
AuthenLoRA: Entangling Stylization with Imperceptible Watermarks for Copyright-Secure LoRA Adapters
von: Shi, Fangming, et al.
Veröffentlicht: (2025)
von: Shi, Fangming, et al.
Veröffentlicht: (2025)
LoRACode: LoRA Adapters for Code Embeddings
von: Chaturvedi, Saumya, et al.
Veröffentlicht: (2025)
von: Chaturvedi, Saumya, et al.
Veröffentlicht: (2025)
Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
von: Shen, Yixin
Veröffentlicht: (2025)
von: Shen, Yixin
Veröffentlicht: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
DragLoRA: Online Optimization of LoRA Adapters for Drag-based Image Editing in Diffusion Model
von: Xia, Siwei, et al.
Veröffentlicht: (2025)
von: Xia, Siwei, et al.
Veröffentlicht: (2025)
Weight space Detection of Backdoors in LoRA Adapters
von: Merenciano, David Puertolas, et al.
Veröffentlicht: (2026)
von: Merenciano, David Puertolas, et al.
Veröffentlicht: (2026)
LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2026)
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2026)
LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
Effective LoRA Adapter Routing using Task Representations
von: Dhasade, Akash, et al.
Veröffentlicht: (2026)
von: Dhasade, Akash, et al.
Veröffentlicht: (2026)
SLAD : Shared LoRA Adapters for Task Specific Distillation
von: Bensaid, Reda, et al.
Veröffentlicht: (2026)
von: Bensaid, Reda, et al.
Veröffentlicht: (2026)
Improving SAM for Camouflaged Object Detection via Dual Stream Adapters
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
Cross-LoRA: A Data-Free LoRA Transfer Framework across Heterogeneous LLMs
von: Xia, Feifan, et al.
Veröffentlicht: (2025)
von: Xia, Feifan, et al.
Veröffentlicht: (2025)
PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRA
von: Wang, Sheng, et al.
Veröffentlicht: (2024)
von: Wang, Sheng, et al.
Veröffentlicht: (2024)
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA Serving
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
von: Li, Allison, et al.
Veröffentlicht: (2025)
von: Li, Allison, et al.
Veröffentlicht: (2025)
SHE-LoRA: Selective Homomorphic Encryption for Federated Tuning with Heterogeneous LoRA
von: Liu, Jianmin, et al.
Veröffentlicht: (2025)
von: Liu, Jianmin, et al.
Veröffentlicht: (2025)
LoRA-PAR: A Flexible Dual-System LoRA Partitioning Approach to Efficient LLM Fine-Tuning
von: Huang, Yining, et al.
Veröffentlicht: (2025)
von: Huang, Yining, et al.
Veröffentlicht: (2025)
BoostLoRA: Growing Effective Rank by Boosting Adapters
von: Anantha, Raviteja, et al.
Veröffentlicht: (2026)
von: Anantha, Raviteja, et al.
Veröffentlicht: (2026)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024)
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024)
Vision as LoRA
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters
von: Ahn, Beomjin, et al.
Veröffentlicht: (2026)
von: Ahn, Beomjin, et al.
Veröffentlicht: (2026)
HypeLoRA: Hyper-Network-Generated LoRA Adapters for Calibrated Language Model Fine-Tuning
von: Trojan, Bartosz, et al.
Veröffentlicht: (2026)
von: Trojan, Bartosz, et al.
Veröffentlicht: (2026)
LoRA Meets Dropout under a Unified Framework
von: Wang, Sheng, et al.
Veröffentlicht: (2024)
von: Wang, Sheng, et al.
Veröffentlicht: (2024)
Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment
von: Fan, Chenghao, et al.
Veröffentlicht: (2025)
von: Fan, Chenghao, et al.
Veröffentlicht: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
Boosting Robust AIGI Detection with LoRA-based Pairwise Training
von: Xia, Ruiyang, et al.
Veröffentlicht: (2026)
von: Xia, Ruiyang, et al.
Veröffentlicht: (2026)
Improving Fisher Information Estimation and Efficiency for LoRA-based LLM Unlearning
von: Kim, Yejin, et al.
Veröffentlicht: (2025)
von: Kim, Yejin, et al.
Veröffentlicht: (2025)
Parametric Retrieval-Augmented Generation using Latent Routing of LoRA Adapters
von: Su, Zhan, et al.
Veröffentlicht: (2025)
von: Su, Zhan, et al.
Veröffentlicht: (2025)
Run LoRA Run: Faster and Lighter LoRA Implementations
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023)
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023)
FREE-Switch: Frequency-based Dynamic LoRA Switch for Style Transfer
von: Zheng, Shenghe, et al.
Veröffentlicht: (2026)
von: Zheng, Shenghe, et al.
Veröffentlicht: (2026)
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
von: Li, Wenbing, et al.
Veröffentlicht: (2025)
von: Li, Wenbing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Empower Vision Applications with LoRA LMM
von: Mi, Liang, et al.
Veröffentlicht: (2024) -
SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
von: Kong, Rui, et al.
Veröffentlicht: (2023) -
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023) -
AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization
von: Li, Qiyang, et al.
Veröffentlicht: (2026) -
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs
von: Ye, Zhengmao, et al.
Veröffentlicht: (2023)