Enregistré dans:
| Auteurs principaux: | Zhan, Huiyou, Zhang, Xuan, Tan, Haisheng, Tian, Han, Yong, Dongping, Zhang, Junyang, Li, Xiang-Yang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2501.09367 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Cloud Native System for LLM Inference Serving
par: Xu, Minxian, et autres
Publié: (2025)
par: Xu, Minxian, et autres
Publié: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
par: Cao, Jiahe, et autres
Publié: (2026)
par: Cao, Jiahe, et autres
Publié: (2026)
Embedding Samples Dispatching for Recommendation Model Training in Edge Environments
par: Li, Guopeng, et autres
Publié: (2025)
par: Li, Guopeng, et autres
Publié: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
par: Wang, Weiye, et autres
Publié: (2026)
par: Wang, Weiye, et autres
Publié: (2026)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
par: Zhang, Yida, et autres
Publié: (2026)
par: Zhang, Yida, et autres
Publié: (2026)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
par: Jiang, Youhe, et autres
Publié: (2025)
par: Jiang, Youhe, et autres
Publié: (2025)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
par: Nguyen, Thanh-Tung, et autres
Publié: (2025)
par: Nguyen, Thanh-Tung, et autres
Publié: (2025)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
par: Lou, Chiheng, et autres
Publié: (2025)
par: Lou, Chiheng, et autres
Publié: (2025)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
par: Yu, Shibo, et autres
Publié: (2025)
par: Yu, Shibo, et autres
Publié: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
par: Han, Yunhe, et autres
Publié: (2026)
par: Han, Yunhe, et autres
Publié: (2026)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
par: Zhang, Mingjin, et autres
Publié: (2024)
par: Zhang, Mingjin, et autres
Publié: (2024)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
par: Jaiswal, Shashwat, et autres
Publié: (2025)
par: Jaiswal, Shashwat, et autres
Publié: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
par: Chow, Will
Publié: (2025)
par: Chow, Will
Publié: (2025)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
par: Dong, Jiangwen, et autres
Publié: (2025)
par: Dong, Jiangwen, et autres
Publié: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
par: Gao, Luyao, et autres
Publié: (2025)
par: Gao, Luyao, et autres
Publié: (2025)
A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
par: Zhang, Yue, et autres
Publié: (2025)
par: Zhang, Yue, et autres
Publié: (2025)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
par: Yang, Zheming, et autres
Publié: (2025)
par: Yang, Zheming, et autres
Publié: (2025)
SynergAI: Edge-to-Cloud Synergy for Architecture-Driven High-Performance Orchestration for AI Inference
par: Stathopoulou, Foteini, et autres
Publié: (2025)
par: Stathopoulou, Foteini, et autres
Publié: (2025)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
par: Wolfrath, Joel, et autres
Publié: (2025)
par: Wolfrath, Joel, et autres
Publié: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
par: Li, Suyi, et autres
Publié: (2024)
par: Li, Suyi, et autres
Publié: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
par: Du, Boxiao, et autres
Publié: (2026)
par: Du, Boxiao, et autres
Publié: (2026)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
par: Khoshsirat, Aria, et autres
Publié: (2024)
par: Khoshsirat, Aria, et autres
Publié: (2024)
LLM-Driven Intent-Based Privacy-Aware Orchestration Across the Cloud-Edge Continuum
par: Su, Zijie, et autres
Publié: (2026)
par: Su, Zijie, et autres
Publié: (2026)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
par: Ghosh, Himel
Publié: (2024)
par: Ghosh, Himel
Publié: (2024)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
par: Duan, Jiangfei, et autres
Publié: (2024)
par: Duan, Jiangfei, et autres
Publié: (2024)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
par: Ruan, Chaoyi, et autres
Publié: (2025)
par: Ruan, Chaoyi, et autres
Publié: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
par: He, Yiyuan, et autres
Publié: (2024)
par: He, Yiyuan, et autres
Publié: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
par: He, Wenhao, et autres
Publié: (2026)
par: He, Wenhao, et autres
Publié: (2026)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
par: He, Xuan, et autres
Publié: (2025)
par: He, Xuan, et autres
Publié: (2025)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
par: Yang, Zheming, et autres
Publié: (2026)
par: Yang, Zheming, et autres
Publié: (2026)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
par: Zhang, Yuning, et autres
Publié: (2025)
par: Zhang, Yuning, et autres
Publié: (2025)
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
par: Zheng, Zihao, et autres
Publié: (2026)
par: Zheng, Zihao, et autres
Publié: (2026)
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
par: Li, Senyao, et autres
Publié: (2025)
par: Li, Senyao, et autres
Publié: (2025)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
par: Yu, Fengze, et autres
Publié: (2025)
par: Yu, Fengze, et autres
Publié: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
par: Wilkins, Grant, et autres
Publié: (2024)
par: Wilkins, Grant, et autres
Publié: (2024)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
par: Sun, Mingyu, et autres
Publié: (2025)
par: Sun, Mingyu, et autres
Publié: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
par: Du, Jiangsu, et autres
Publié: (2025)
par: Du, Jiangsu, et autres
Publié: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
par: Liao, Junhan, et autres
Publié: (2025)
par: Liao, Junhan, et autres
Publié: (2025)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
par: Zhu, Juan, et autres
Publié: (2026)
par: Zhu, Juan, et autres
Publié: (2026)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
par: Arya, Mayank, et autres
Publié: (2025)
par: Arya, Mayank, et autres
Publié: (2025)
Documents similaires
-
Cloud Native System for LLM Inference Serving
par: Xu, Minxian, et autres
Publié: (2025) -
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
par: Cao, Jiahe, et autres
Publié: (2026) -
Embedding Samples Dispatching for Recommendation Model Training in Edge Environments
par: Li, Guopeng, et autres
Publié: (2025) -
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
par: Wang, Weiye, et autres
Publié: (2026) -
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
par: Zhang, Yida, et autres
Publié: (2026)