PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gao, Wei, Sun, Peng, Ustiugov, Dmitrii, Zhang, Tianwei, Wen, Yonggang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
par: Chen, Qiaoling, et autres
Publié: (2025)
par: Chen, Qiaoling, et autres
Publié: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
par: Chen, Wenyan, et autres
Publié: (2026)
par: Chen, Wenyan, et autres
Publié: (2026)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
par: Nie, Chengyi, et autres
Publié: (2024)
par: Nie, Chengyi, et autres
Publié: (2024)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
par: Wang, Qipeng
Publié: (2026)
par: Wang, Qipeng
Publié: (2026)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
par: Chen, Qiaoling, et autres
Publié: (2023)
par: Chen, Qiaoling, et autres
Publié: (2023)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
par: Lai, Ruiqi, et autres
Publié: (2025)
par: Lai, Ruiqi, et autres
Publié: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
par: Chen, Qiaoling, et autres
Publié: (2026)
par: Chen, Qiaoling, et autres
Publié: (2026)
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
par: Wang, Zerui, et autres
Publié: (2025)
par: Wang, Zerui, et autres
Publié: (2025)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
par: Sedlak, Boris, et autres
Publié: (2024)
par: Sedlak, Boris, et autres
Publié: (2024)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
par: Chow, Will
Publié: (2025)
par: Chow, Will
Publié: (2025)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
par: Peng, Haosong, et autres
Publié: (2024)
par: Peng, Haosong, et autres
Publié: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
par: Shen, Haiying, et autres
Publié: (2024)
par: Shen, Haiying, et autres
Publié: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
par: Mo, Zizhao, et autres
Publié: (2026)
par: Mo, Zizhao, et autres
Publié: (2026)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
par: Fu, Yao, et autres
Publié: (2024)
par: Fu, Yao, et autres
Publié: (2024)
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
par: Willemsen, Floris-Jan, et autres
Publié: (2025)
par: Willemsen, Floris-Jan, et autres
Publié: (2025)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
par: Sun, Desen, et autres
Publié: (2025)
par: Sun, Desen, et autres
Publié: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
par: Hu, Jianmin, et autres
Publié: (2025)
par: Hu, Jianmin, et autres
Publié: (2025)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
par: Xuan, Mo, et autres
Publié: (2025)
par: Xuan, Mo, et autres
Publié: (2025)
Multi-Modal Style Transfer-based Prompt Tuning for Efficient Federated Domain Generalization
par: Chen, Yuliang, et autres
Publié: (2026)
par: Chen, Yuliang, et autres
Publié: (2026)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
par: Hu, Kan, et autres
Publié: (2024)
par: Hu, Kan, et autres
Publié: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
par: Hu, Tiancheng, et autres
Publié: (2026)
par: Hu, Tiancheng, et autres
Publié: (2026)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
par: Jeon, Beomyeol, et autres
Publié: (2024)
par: Jeon, Beomyeol, et autres
Publié: (2024)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
par: Chen, Qiaoling, et autres
Publié: (2025)
par: Chen, Qiaoling, et autres
Publié: (2025)
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
par: Kondrashov, Leonid, et autres
Publié: (2025)
par: Kondrashov, Leonid, et autres
Publié: (2025)
TorchGT: A Holistic System for Large-scale Graph Transformer Training
par: Zhang, Meng, et autres
Publié: (2024)
par: Zhang, Meng, et autres
Publié: (2024)
An SLO Driven and Cost-Aware Autoscaling Framework for Kubernetes
par: Punniyamoorthy, Vinoth, et autres
Publié: (2025)
par: Punniyamoorthy, Vinoth, et autres
Publié: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
par: Ma, Chenxiang, et autres
Publié: (2025)
par: Ma, Chenxiang, et autres
Publié: (2025)
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
par: Chen, Qiaoling, et autres
Publié: (2024)
par: Chen, Qiaoling, et autres
Publié: (2024)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
par: Qi, S., et autres
Publié: (2024)
par: Qi, S., et autres
Publié: (2024)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
par: Xu, Ao, et autres
Publié: (2025)
par: Xu, Ao, et autres
Publié: (2025)
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
par: Park, JooYoung, et autres
Publié: (2026)
par: Park, JooYoung, et autres
Publié: (2026)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)
par: Chen, Jiabin, et autres
Publié: (2024)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
par: Huang, Shaoyuan, et autres
Publié: (2026)
par: Huang, Shaoyuan, et autres
Publié: (2026)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
par: Xu, Jiale, et autres
Publié: (2025)
par: Xu, Jiale, et autres
Publié: (2025)
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
par: Li, Luchang, et autres
Publié: (2026)
par: Li, Luchang, et autres
Publié: (2026)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
par: Zhang, Zongshun, et autres
Publié: (2025)
par: Zhang, Zongshun, et autres
Publié: (2025)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
par: Gu, Diandian, et autres
Publié: (2024)
par: Gu, Diandian, et autres
Publié: (2024)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
par: Hu, Kan, et autres
Publié: (2024)
par: Hu, Kan, et autres
Publié: (2024)
Documents similaires
-
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
par: Chen, Qiaoling, et autres
Publié: (2025) -
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
par: Chen, Wenyan, et autres
Publié: (2026) -
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
par: Cheng, Ke, et autres
Publié: (2024) -
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
par: Nie, Chengyi, et autres
Publié: (2024) -
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
par: Wang, Qipeng
Publié: (2026)