LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jeon, Hyesung, Ha, Hyeongju, Kim, Jae-Joon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
von: Wang, Shao, et al.
Veröffentlicht: (2026)
von: Wang, Shao, et al.
Veröffentlicht: (2026)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
von: Li, Allison, et al.
Veröffentlicht: (2025)
von: Li, Allison, et al.
Veröffentlicht: (2025)
tLoRA: Efficient Multi-LoRA Training with Elastic Shared Super-Models
von: Li, Kevin, et al.
Veröffentlicht: (2026)
von: Li, Kevin, et al.
Veröffentlicht: (2026)
R-LoRA: Randomized Multi-Head LoRA for Efficient Multi-Task Learning
von: Liu, Jinda, et al.
Veröffentlicht: (2025)
von: Liu, Jinda, et al.
Veröffentlicht: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models
von: Chen, Guanduo, et al.
Veröffentlicht: (2025)
von: Chen, Guanduo, et al.
Veröffentlicht: (2025)
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
von: Jeon, Hyesung, et al.
Veröffentlicht: (2024)
von: Jeon, Hyesung, et al.
Veröffentlicht: (2024)
LoRA-PAR: A Flexible Dual-System LoRA Partitioning Approach to Efficient LLM Fine-Tuning
von: Huang, Yining, et al.
Veröffentlicht: (2025)
von: Huang, Yining, et al.
Veröffentlicht: (2025)
LoRA-drop: Efficient LoRA Parameter Pruning based on Output Evaluation
von: Zhou, Hongyun, et al.
Veröffentlicht: (2024)
von: Zhou, Hongyun, et al.
Veröffentlicht: (2024)
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving
von: Li, Shaoang, et al.
Veröffentlicht: (2026)
von: Li, Shaoang, et al.
Veröffentlicht: (2026)
Run LoRA Run: Faster and Lighter LoRA Implementations
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023)
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023)
Improving Fisher Information Estimation and Efficiency for LoRA-based LLM Unlearning
von: Kim, Yejin, et al.
Veröffentlicht: (2025)
von: Kim, Yejin, et al.
Veröffentlicht: (2025)
PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRA
von: Wang, Sheng, et al.
Veröffentlicht: (2024)
von: Wang, Sheng, et al.
Veröffentlicht: (2024)
LoRA Diffusion: Zero-Shot LoRA Synthesis for Diffusion Model Personalization
von: Smith, Ethan, et al.
Veröffentlicht: (2024)
von: Smith, Ethan, et al.
Veröffentlicht: (2024)
LoRA-MME: Multi-Model Ensemble of LoRA-Tuned Encoders for Code Comment Classification
von: Haider, Md Akib, et al.
Veröffentlicht: (2026)
von: Haider, Md Akib, et al.
Veröffentlicht: (2026)
PLoRA: Efficient LoRA Hyperparameter Tuning for Large Models
von: Yan, Minghao, et al.
Veröffentlicht: (2025)
von: Yan, Minghao, et al.
Veröffentlicht: (2025)
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
von: Bae, Sangmin, et al.
Veröffentlicht: (2024)
von: Bae, Sangmin, et al.
Veröffentlicht: (2024)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
von: Patel, Ishan, et al.
Veröffentlicht: (2026)
von: Patel, Ishan, et al.
Veröffentlicht: (2026)
SC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA
von: Luo, Minrui, et al.
Veröffentlicht: (2025)
von: Luo, Minrui, et al.
Veröffentlicht: (2025)
Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA Serving
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
KD-LoRA: A Hybrid Approach to Efficient Fine-Tuning with LoRA and Knowledge Distillation
von: Azimi, Rambod, et al.
Veröffentlicht: (2024)
von: Azimi, Rambod, et al.
Veröffentlicht: (2024)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
GeLoRA: Geometric Adaptive Ranks For Efficient LoRA Fine-tuning
von: Ed-dib, Abdessalam, et al.
Veröffentlicht: (2024)
von: Ed-dib, Abdessalam, et al.
Veröffentlicht: (2024)
LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)
Cross-LoRA: A Data-Free LoRA Transfer Framework across Heterogeneous LLMs
von: Xia, Feifan, et al.
Veröffentlicht: (2025)
von: Xia, Feifan, et al.
Veröffentlicht: (2025)
Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
von: Liu, Yezi, et al.
Veröffentlicht: (2025)
von: Liu, Yezi, et al.
Veröffentlicht: (2025)
FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance
von: Chung, Hyunsuk, et al.
Veröffentlicht: (2026)
von: Chung, Hyunsuk, et al.
Veröffentlicht: (2026)
LoRA on the Go: Instance-level Dynamic LoRA Selection and Merging
von: Lee, Seungeon, et al.
Veröffentlicht: (2025)
von: Lee, Seungeon, et al.
Veröffentlicht: (2025)
Latent Space Factorization in LoRA
von: Kumar, Shashi, et al.
Veröffentlicht: (2025)
von: Kumar, Shashi, et al.
Veröffentlicht: (2025)
Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
von: Shen, Yixin
Veröffentlicht: (2025)
von: Shen, Yixin
Veröffentlicht: (2025)
AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
von: Li, Jiacheng, et al.
Veröffentlicht: (2025)
von: Li, Jiacheng, et al.
Veröffentlicht: (2025)
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
von: Li, Wenbing, et al.
Veröffentlicht: (2025)
von: Li, Wenbing, et al.
Veröffentlicht: (2025)
Riemannian Optimization for LoRA on the Stiefel Manifold
von: Park, Juneyoung, et al.
Veröffentlicht: (2025)
von: Park, Juneyoung, et al.
Veröffentlicht: (2025)
Shared LoRA Subspaces for almost Strict Continual Learning
von: Kaushik, Prakhar, et al.
Veröffentlicht: (2026)
von: Kaushik, Prakhar, et al.
Veröffentlicht: (2026)
Tensorized Clustered LoRA Merging for Multi-Task Interference
von: Su, Zhan, et al.
Veröffentlicht: (2025)
von: Su, Zhan, et al.
Veröffentlicht: (2025)
LoRA-FAIR: Federated LoRA Fine-Tuning with Aggregation and Initialization Refinement
von: Bian, Jieming, et al.
Veröffentlicht: (2024)
von: Bian, Jieming, et al.
Veröffentlicht: (2024)
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
von: Yen, Jui-Nan, et al.
Veröffentlicht: (2024)
von: Yen, Jui-Nan, et al.
Veröffentlicht: (2024)
Mixture of LoRA Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
MoRA: LoRA Guided Multi-Modal Disease Diagnosis with Missing Modality
von: Shi, Zhiyi, et al.
Veröffentlicht: (2024)
von: Shi, Zhiyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025) -
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
von: Wang, Shao, et al.
Veröffentlicht: (2026) -
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
von: Li, Allison, et al.
Veröffentlicht: (2025) -
tLoRA: Efficient Multi-LoRA Training with Elastic Shared Super-Models
von: Li, Kevin, et al.
Veröffentlicht: (2026) -
R-LoRA: Randomized Multi-Head LoRA for Efficient Multi-Task Learning
von: Liu, Jinda, et al.
Veröffentlicht: (2025)