Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Peng, Haosong, Zhan, Yufeng, Li, Peng, Xia, Yuanqing |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Egret: Reinforcement Mechanism for Sequential Computation Offloading in Edge Computing
par: Peng, Haosong, et autres
Publié: (2024)
par: Peng, Haosong, et autres
Publié: (2024)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)
par: Chen, Jiabin, et autres
Publié: (2024)
Radiant: Large-scale 3D Gaussian Rendering based on Hierarchical Framework
par: Peng, Haosong, et autres
Publié: (2024)
par: Peng, Haosong, et autres
Publié: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
par: Gu, Jianfeng, et autres
Publié: (2025)
par: Gu, Jianfeng, et autres
Publié: (2025)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
par: Yu, Minchen, et autres
Publié: (2025)
par: Yu, Minchen, et autres
Publié: (2025)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
par: Qi, S., et autres
Publié: (2024)
par: Qi, S., et autres
Publié: (2024)
Affinity-aware Serverless Function Scheduling
par: De Palma, Giuseppe, et autres
Publié: (2024)
par: De Palma, Giuseppe, et autres
Publié: (2024)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
par: Xuan, Mo, et autres
Publié: (2025)
par: Xuan, Mo, et autres
Publié: (2025)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
par: Gao, Wei, et autres
Publié: (2026)
par: Gao, Wei, et autres
Publié: (2026)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
par: Zhu, Wenbin, et autres
Publié: (2025)
par: Zhu, Wenbin, et autres
Publié: (2025)
Dependency-aware Resource Allocation for Serverless Functions at the Edge
par: Baresi, Luciano, et autres
Publié: (2023)
par: Baresi, Luciano, et autres
Publié: (2023)
GreenWhisk: Emission-Aware Computing for Serverless Platform
par: Serenari, Jayden, et autres
Publié: (2024)
par: Serenari, Jayden, et autres
Publié: (2024)
Komet: A Serverless Platform for Low-Earth Orbit Edge Services
par: Pfandzelter, Tobias, et autres
Publié: (2024)
par: Pfandzelter, Tobias, et autres
Publié: (2024)
Saarthi: An End-to-End Intelligent Platform for Optimising Distributed Serverless Workloads
par: Agarwal, Siddharth, et autres
Publié: (2025)
par: Agarwal, Siddharth, et autres
Publié: (2025)
New Kids: An Architecture and Performance Investigation of Second-Generation Serverless Platforms
par: Schirmer, Trever, et autres
Publié: (2026)
par: Schirmer, Trever, et autres
Publié: (2026)
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
ProFaaStinate: Delaying Serverless Function Calls to Optimize Platform Performance
par: Schirmer, Trever, et autres
Publié: (2023)
par: Schirmer, Trever, et autres
Publié: (2023)
Towards Fast Setup and High Throughput of GPU Serverless Computing
par: Zhao, Han, et autres
Publié: (2024)
par: Zhao, Han, et autres
Publié: (2024)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
par: Duan, Jiaang, et autres
Publié: (2024)
par: Duan, Jiaang, et autres
Publié: (2024)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
par: Hui, Xinning, et autres
Publié: (2024)
par: Hui, Xinning, et autres
Publié: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
par: Hu, Tiancheng, et autres
Publié: (2026)
par: Hu, Tiancheng, et autres
Publié: (2026)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
par: Ekelund, Jonah, et autres
Publié: (2025)
par: Ekelund, Jonah, et autres
Publié: (2025)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
par: Pang, Bowen, et autres
Publié: (2025)
par: Pang, Bowen, et autres
Publié: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
par: Ma, Chenxiang, et autres
Publié: (2025)
par: Ma, Chenxiang, et autres
Publié: (2025)
Hydra: Virtualized Multi-Language Runtime for High-Density Serverless Platforms
par: Ivanenko, Serhii, et autres
Publié: (2022)
par: Ivanenko, Serhii, et autres
Publié: (2022)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
par: Nie, Chengyi, et autres
Publié: (2024)
par: Nie, Chengyi, et autres
Publié: (2024)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
par: Sedlak, Boris, et autres
Publié: (2024)
par: Sedlak, Boris, et autres
Publié: (2024)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
par: Lin, Yanying, et autres
Publié: (2025)
par: Lin, Yanying, et autres
Publié: (2025)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
par: Chen, Qiaoling, et autres
Publié: (2026)
par: Chen, Qiaoling, et autres
Publié: (2026)
Cppless: Single-Source and High-Performance Serverless Programming in C++
par: Copik, Marcin, et autres
Publié: (2024)
par: Copik, Marcin, et autres
Publié: (2024)
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
par: Bai, Fengyao, et autres
Publié: (2026)
par: Bai, Fengyao, et autres
Publié: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
par: Wang, Qipeng
Publié: (2026)
par: Wang, Qipeng
Publié: (2026)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
par: Chow, Will
Publié: (2025)
par: Chow, Will
Publié: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
par: Barrak, Amine, et autres
Publié: (2025)
par: Barrak, Amine, et autres
Publié: (2025)
Are Unikernels Ready for Serverless on the Edge?
par: Moebius, Felix, et autres
Publié: (2024)
par: Moebius, Felix, et autres
Publié: (2024)
Jiagu: Optimizing Serverless Computing Resource Utilization with Harmonized Efficiency and Practicability
par: Liu, Qingyuan, et autres
Publié: (2024)
par: Liu, Qingyuan, et autres
Publié: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
par: Shen, Haiying, et autres
Publié: (2024)
par: Shen, Haiying, et autres
Publié: (2024)
Zenix: Efficient Execution of Bulky Serverless Applications
par: Guo, Zhiyuan, et autres
Publié: (2022)
par: Guo, Zhiyuan, et autres
Publié: (2022)
Documents similaires
-
Egret: Reinforcement Mechanism for Sequential Computation Offloading in Edge Computing
par: Peng, Haosong, et autres
Publié: (2024) -
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024) -
Radiant: Large-scale 3D Gaussian Rendering based on Hierarchical Framework
par: Peng, Haosong, et autres
Publié: (2024) -
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
par: Gu, Jianfeng, et autres
Publié: (2025) -
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
par: Yu, Minchen, et autres
Publié: (2025)