Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ni, Yinan, Yang, Xiao, Tang, Yuqi, Qiu, Zhimin, Wang, Chen, Yuan, Tingzhou |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Graph-Structured Deep Learning Framework for Multi-task Contention Identification with High-dimensional Metrics
par: Yang, Xiao, et autres
Publié: (2026)
par: Yang, Xiao, et autres
Publié: (2026)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
par: Sui, Yifan, et autres
Publié: (2025)
par: Sui, Yifan, et autres
Publié: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
par: Li, Suyi, et autres
Publié: (2024)
par: Li, Suyi, et autres
Publié: (2024)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
par: Chen, Hongyu, et autres
Publié: (2026)
par: Chen, Hongyu, et autres
Publié: (2026)
SHE-LoRA: Selective Homomorphic Encryption for Federated Tuning with Heterogeneous LoRA
par: Liu, Jianmin, et autres
Publié: (2025)
par: Liu, Jianmin, et autres
Publié: (2025)
FDLoRA: Personalized Federated Learning of Large Language Model via Dual LoRA Tuning
par: QI, Jiaxing, et autres
Publié: (2024)
par: QI, Jiaxing, et autres
Publié: (2024)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
par: Sheng, Ying, et autres
Publié: (2023)
par: Sheng, Ying, et autres
Publié: (2023)
Federated LoRA with Sparse Communication
par: Kuo, Kevin, et autres
Publié: (2024)
par: Kuo, Kevin, et autres
Publié: (2024)
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
par: Zhu, Zhanda, et autres
Publié: (2025)
par: Zhu, Zhanda, et autres
Publié: (2025)
Stabilizing Decentralized Federated Fine-Tuning via Topology-Aware Alternating LoRA
par: Wang, Xiaoyu, et autres
Publié: (2026)
par: Wang, Xiaoyu, et autres
Publié: (2026)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
par: Jaiswal, Shashwat, et autres
Publié: (2025)
par: Jaiswal, Shashwat, et autres
Publié: (2025)
LoRA-C: Parameter-Efficient Fine-Tuning of Robust CNN for IoT Devices
par: Ding, Chuntao, et autres
Publié: (2024)
par: Ding, Chuntao, et autres
Publié: (2024)
Taming Cold Starts: Proactive Serverless Scheduling with Model Predictive Control
par: Nguyen, Chanh, et autres
Publié: (2025)
par: Nguyen, Chanh, et autres
Publié: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
par: Xu, Chuhao, et autres
Publié: (2025)
par: Xu, Chuhao, et autres
Publié: (2025)
pFedLoRA: Model-Heterogeneous Personalized Federated Learning with LoRA Tuning
par: Yi, Liping, et autres
Publié: (2023)
par: Yi, Liping, et autres
Publié: (2023)
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
par: Zuo, Jingwei, et autres
Publié: (2026)
par: Zuo, Jingwei, et autres
Publié: (2026)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
par: Chang, Zihan, et autres
Publié: (2024)
par: Chang, Zihan, et autres
Publié: (2024)
FedQuad: Adaptive Layer-wise LoRA Deployment and Activation Quantization for Federated Fine-Tuning
par: Li, Rukuo, et autres
Publié: (2025)
par: Li, Rukuo, et autres
Publié: (2025)
Robust Federated Finetuning of Foundation Models via Alternating Minimization of LoRA
par: Chen, Shuangyi, et autres
Publié: (2024)
par: Chen, Shuangyi, et autres
Publié: (2024)
AutoRank: MCDA Based Rank Personalization for LoRA-Enabled Distributed Learning
par: Chen, Shuaijun, et autres
Publié: (2024)
par: Chen, Shuaijun, et autres
Publié: (2024)
Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation Models
par: Cho, Yae Jee, et autres
Publié: (2024)
par: Cho, Yae Jee, et autres
Publié: (2024)
RBLA: Rank-Based-LoRA-Aggregation for Fine-tuning Heterogeneous Models in FLaaS
par: Chen, Shuaijun, et autres
Publié: (2024)
par: Chen, Shuaijun, et autres
Publié: (2024)
AAPA: An Archetype-Aware Predictive Autoscaler with Uncertainty Quantification for Serverless Workloads on Kubernetes
par: Zhang, Guilin, et autres
Publié: (2025)
par: Zhang, Guilin, et autres
Publié: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
par: Yu, Minchen, et autres
Publié: (2023)
par: Yu, Minchen, et autres
Publié: (2023)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
par: Yu, Minchen, et autres
Publié: (2025)
par: Yu, Minchen, et autres
Publié: (2025)
FedRPCA: Enhancing Federated LoRA Aggregation Using Robust PCA
par: Jhunjhunwala, Divyansh, et autres
Publié: (2025)
par: Jhunjhunwala, Divyansh, et autres
Publié: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)
par: Chen, Jiabin, et autres
Publié: (2024)
Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?
par: Wang, Yatong, et autres
Publié: (2026)
par: Wang, Yatong, et autres
Publié: (2026)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
par: Lin, Yanying, et autres
Publié: (2025)
par: Lin, Yanying, et autres
Publié: (2025)
GreenWhisk: Emission-Aware Computing for Serverless Platform
par: Serenari, Jayden, et autres
Publié: (2024)
par: Serenari, Jayden, et autres
Publié: (2024)
Energy Efficient Scheduling for Serverless Systems
par: Tsenos, Michail, et autres
Publié: (2024)
par: Tsenos, Michail, et autres
Publié: (2024)
Improving LoRA in Privacy-preserving Federated Learning
par: Sun, Youbang, et autres
Publié: (2024)
par: Sun, Youbang, et autres
Publié: (2024)
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
par: Duan, Jiaang, et autres
Publié: (2024)
par: Duan, Jiaang, et autres
Publié: (2024)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
par: Wu, Z., et autres
Publié: (2025)
par: Wu, Z., et autres
Publié: (2025)
Fed-pilot: Optimizing LoRA Allocation for Efficient Federated Fine-Tuning with Heterogeneous Clients
par: Zhang, Zikai, et autres
Publié: (2024)
par: Zhang, Zikai, et autres
Publié: (2024)
LoRA-based Parameter-Efficient LLMs for Continuous Learning in Edge-based Malware Detection
par: Rondanini, Christian, et autres
Publié: (2026)
par: Rondanini, Christian, et autres
Publié: (2026)
Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
par: Li, Rui, et autres
Publié: (2025)
par: Li, Rui, et autres
Publié: (2025)
PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
par: Zhao, Zhixin, et autres
Publié: (2026)
par: Zhao, Zhixin, et autres
Publié: (2026)
Spatiotemporal Traffic Prediction in Distributed Backend Systems via Graph Neural Networks
par: Qiu, Zhimin, et autres
Publié: (2025)
par: Qiu, Zhimin, et autres
Publié: (2025)
EcoLife: Carbon-Aware Serverless Function Scheduling for Sustainable Computing
par: Jiang, Yankai, et autres
Publié: (2024)
par: Jiang, Yankai, et autres
Publié: (2024)
Documents similaires
-
Graph-Structured Deep Learning Framework for Multi-task Contention Identification with High-dimensional Metrics
par: Yang, Xiao, et autres
Publié: (2026) -
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
par: Sui, Yifan, et autres
Publié: (2025) -
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
par: Li, Suyi, et autres
Publié: (2024) -
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
par: Chen, Hongyu, et autres
Publié: (2026) -
SHE-LoRA: Selective Homomorphic Encryption for Federated Tuning with Heterogeneous LoRA
par: Liu, Jianmin, et autres
Publié: (2025)