Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Jaiswal, Shashwat, Arun, Shrikara, Parayil, Anjaly, Mallick, Ankur, Mastorakis, Spyros, Khare, Alind, Alverti, Chloi, Amant, Renee St, Bansal, Chetan, Rühle, Victor, Torrellas, Josep |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
Ensuring Fair LLM Serving Amid Diverse Applications
by: Khan, Redwan Ibne Seraj, et al.
Published: (2024)
by: Khan, Redwan Ibne Seraj, et al.
Published: (2024)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
by: Biswas, Anish, et al.
Published: (2026)
by: Biswas, Anish, et al.
Published: (2026)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
by: Sheng, Ying, et al.
Published: (2023)
by: Sheng, Ying, et al.
Published: (2023)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024)
by: Jain, Kunal, et al.
Published: (2024)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)
by: Sanovar, Rya, et al.
Published: (2024)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
Intelligent Monitoring Framework for Cloud Services: A Data-Driven Approach
by: Srinivas, Pooja, et al.
Published: (2024)
by: Srinivas, Pooja, et al.
Published: (2024)
A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring
by: Bastos, Anson, et al.
Published: (2026)
by: Bastos, Anson, et al.
Published: (2026)
AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
by: Barke, Shraddha, et al.
Published: (2026)
by: Barke, Shraddha, et al.
Published: (2026)
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
by: Couturier, Camille, et al.
Published: (2025)
by: Couturier, Camille, et al.
Published: (2025)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
by: Qiu, Haoran, et al.
Published: (2025)
by: Qiu, Haoran, et al.
Published: (2025)
Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems
by: Hussain, Fiza, et al.
Published: (2025)
by: Hussain, Fiza, et al.
Published: (2025)
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)
by: Parayil, Anjaly, et al.
Published: (2024)
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving
by: Li, Shaoang, et al.
Published: (2026)
by: Li, Shaoang, et al.
Published: (2026)
X-lifecycle Learning for Cloud Incident Management using LLMs
by: Goel, Drishti, et al.
Published: (2024)
by: Goel, Drishti, et al.
Published: (2024)
Gaussian beams and inverse problems for connections at high fixed frequency
by: St-Amant, Simon
Published: (2024)
by: St-Amant, Simon
Published: (2024)
The Gulf States Marine Fisheries Commission and the Menhaden Fishery
by: St. Amant, L.S.
Published: (1974)
by: St. Amant, L.S.
Published: (1974)
Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
by: Shen, Yixin
Published: (2025)
by: Shen, Yixin
Published: (2025)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
by: Li, Allison, et al.
Published: (2025)
by: Li, Allison, et al.
Published: (2025)
Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA Serving
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
LoRACode: LoRA Adapters for Code Embeddings
by: Chaturvedi, Saumya, et al.
Published: (2025)
by: Chaturvedi, Saumya, et al.
Published: (2025)
Weight space Detection of Backdoors in LoRA Adapters
by: Merenciano, David Puertolas, et al.
Published: (2026)
by: Merenciano, David Puertolas, et al.
Published: (2026)
LLMs for Generation of Architectural Components: An Exploratory Empirical Study in the Serverless World
by: Arun, Shrikara, et al.
Published: (2025)
by: Arun, Shrikara, et al.
Published: (2025)
Towards CXL Resilience to CPU Failures
by: Psistakis, Antonis, et al.
Published: (2026)
by: Psistakis, Antonis, et al.
Published: (2026)
LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
Effective LoRA Adapter Routing using Task Representations
by: Dhasade, Akash, et al.
Published: (2026)
by: Dhasade, Akash, et al.
Published: (2026)
SLAD : Shared LoRA Adapters for Task Specific Distillation
by: Bensaid, Reda, et al.
Published: (2026)
by: Bensaid, Reda, et al.
Published: (2026)
Characterizing Stability in Many-to-One Matching with Non-Responsive Couples
by: Khare, Shashwat, et al.
Published: (2025)
by: Khare, Shashwat, et al.
Published: (2025)
AuthenLoRA: Entangling Stylization with Imperceptible Watermarks for Copyright-Secure LoRA Adapters
by: Shi, Fangming, et al.
Published: (2025)
by: Shi, Fangming, et al.
Published: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
by: Chen, Hongyu, et al.
Published: (2026)
by: Chen, Hongyu, et al.
Published: (2026)
Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC
by: Soni, Aditya, et al.
Published: (2024)
by: Soni, Aditya, et al.
Published: (2024)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
by: Zhang, Hang, et al.
Published: (2025)
by: Zhang, Hang, et al.
Published: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
by: Li, Suyi, et al.
Published: (2024)
by: Li, Suyi, et al.
Published: (2024)
Towards Architecting Sustainable MLOps: A Self-Adaptation Approach
by: Bhatt, Hiya, et al.
Published: (2024)
by: Bhatt, Hiya, et al.
Published: (2024)
Parametric Retrieval-Augmented Generation using Latent Routing of LoRA Adapters
by: Su, Zhan, et al.
Published: (2025)
by: Su, Zhan, et al.
Published: (2025)
Management guidelines for predicting brown shrimp, Penaeus aztecus, production in Louisiana
by: Ford, T.B., et al.
Published: (1971)
by: Ford, T.B., et al.
Published: (1971)
HypeLoRA: Hyper-Network-Generated LoRA Adapters for Calibrated Language Model Fine-Tuning
by: Trojan, Bartosz, et al.
Published: (2026)
by: Trojan, Bartosz, et al.
Published: (2026)
DragLoRA: Online Optimization of LoRA Adapters for Drag-based Image Editing in Diffusion Model
by: Xia, Siwei, et al.
Published: (2025)
by: Xia, Siwei, et al.
Published: (2025)
Similar Items
-
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025) -
Ensuring Fair LLM Serving Amid Diverse Applications
by: Khan, Redwan Ibne Seraj, et al.
Published: (2024) -
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
by: Biswas, Anish, et al.
Published: (2026) -
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
by: Sheng, Ying, et al.
Published: (2023) -
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024)