TimelyLLM: Segmented LLM Serving System for Time-sensitive Robotic Applications
Fuente:
arXiv
Guardado en:
| Autores principales: | Ling, Neiwen, Chen, Guojun, Zhong, Lin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
por: Agrawal, Amey, et al.
Publicado: (2026)
por: Agrawal, Amey, et al.
Publicado: (2026)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
por: Lou, Chiheng, et al.
Publicado: (2025)
por: Lou, Chiheng, et al.
Publicado: (2025)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
por: Wu, Bingyang, et al.
Publicado: (2025)
por: Wu, Bingyang, et al.
Publicado: (2025)
Locality-aware Fair Scheduling in LLM Serving
por: Cao, Shiyi, et al.
Publicado: (2025)
por: Cao, Shiyi, et al.
Publicado: (2025)
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
por: Liu, Yifei, et al.
Publicado: (2025)
por: Liu, Yifei, et al.
Publicado: (2025)
Floe: Federated Specialization for Real-Time LLM-SLM Inference
por: Tian, Chunlin, et al.
Publicado: (2026)
por: Tian, Chunlin, et al.
Publicado: (2026)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
por: Qiao, Yifan, et al.
Publicado: (2024)
por: Qiao, Yifan, et al.
Publicado: (2024)
Preble: Efficient Distributed Prompt Scheduling for LLM Serving
por: Srivatsa, Vikranth, et al.
Publicado: (2024)
por: Srivatsa, Vikranth, et al.
Publicado: (2024)
Collaborative Speculative Inference for Efficient LLM Inference Serving
por: Gao, Luyao, et al.
Publicado: (2025)
por: Gao, Luyao, et al.
Publicado: (2025)
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
por: Yao, Yuhang, et al.
Publicado: (2024)
por: Yao, Yuhang, et al.
Publicado: (2024)
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
por: Liang, Yunkai, et al.
Publicado: (2025)
por: Liang, Yunkai, et al.
Publicado: (2025)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
por: Agrawal, Amey, et al.
Publicado: (2024)
por: Agrawal, Amey, et al.
Publicado: (2024)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
por: Bin, Kyungmin, et al.
Publicado: (2025)
por: Bin, Kyungmin, et al.
Publicado: (2025)
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
por: Bai, Xu, et al.
Publicado: (2026)
por: Bai, Xu, et al.
Publicado: (2026)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
por: Xu, Jiale, et al.
Publicado: (2024)
por: Xu, Jiale, et al.
Publicado: (2024)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
por: Shi, Xiaoxiang, et al.
Publicado: (2025)
por: Shi, Xiaoxiang, et al.
Publicado: (2025)
On Evaluating Performance of LLM Inference Serving Systems
por: Agrawal, Amey, et al.
Publicado: (2025)
por: Agrawal, Amey, et al.
Publicado: (2025)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
por: Su, Zhaoyuan, et al.
Publicado: (2025)
por: Su, Zhaoyuan, et al.
Publicado: (2025)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
por: Ghosh, Himel
Publicado: (2024)
por: Ghosh, Himel
Publicado: (2024)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
por: Kim, Kihyun, et al.
Publicado: (2025)
por: Kim, Kihyun, et al.
Publicado: (2025)
Kairos: A Scalable Serving System for Physical AI
por: Dai, Yinwei, et al.
Publicado: (2026)
por: Dai, Yinwei, et al.
Publicado: (2026)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
por: Zhao, Juntao, et al.
Publicado: (2024)
por: Zhao, Juntao, et al.
Publicado: (2024)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
por: Woo, Sunghyeon, et al.
Publicado: (2026)
por: Woo, Sunghyeon, et al.
Publicado: (2026)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
por: Liu, Xueshen, et al.
Publicado: (2026)
por: Liu, Xueshen, et al.
Publicado: (2026)
HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
por: Sun, Ting, et al.
Publicado: (2025)
por: Sun, Ting, et al.
Publicado: (2025)
COHORT: Hybrid RL for Collaborative Large DNN Inference on Multi-Robot Systems Under Real-Time Constraints
por: Anwar, Mohammad Saeid, et al.
Publicado: (2026)
por: Anwar, Mohammad Saeid, et al.
Publicado: (2026)
Gradient-based Trajectory Optimization with Parallelized Differentiable Traffic Simulation
por: Son, Sanghyun, et al.
Publicado: (2024)
por: Son, Sanghyun, et al.
Publicado: (2024)
Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
por: Xu, Yechen, et al.
Publicado: (2024)
por: Xu, Yechen, et al.
Publicado: (2024)
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
por: Yu, Jiahuan, et al.
Publicado: (2025)
por: Yu, Jiahuan, et al.
Publicado: (2025)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
por: Oliaro, Gabriele, et al.
Publicado: (2024)
por: Oliaro, Gabriele, et al.
Publicado: (2024)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
por: Ye, Zihao, et al.
Publicado: (2025)
por: Ye, Zihao, et al.
Publicado: (2025)
RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models
por: Zheng, Zihao, et al.
Publicado: (2026)
por: Zheng, Zihao, et al.
Publicado: (2026)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
por: Wu, Bingyang, et al.
Publicado: (2024)
por: Wu, Bingyang, et al.
Publicado: (2024)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
por: Chen, Siyuan, et al.
Publicado: (2025)
por: Chen, Siyuan, et al.
Publicado: (2025)
Communication Resources Constrained Hierarchical Federated Learning for End-to-End Autonomous Driving
por: Kou, Wei-Bin, et al.
Publicado: (2023)
por: Kou, Wei-Bin, et al.
Publicado: (2023)
Niyama : Breaking the Silos of LLM Inference Serving
por: Goel, Kanishk, et al.
Publicado: (2025)
por: Goel, Kanishk, et al.
Publicado: (2025)
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
por: Hong, Ke, et al.
Publicado: (2025)
por: Hong, Ke, et al.
Publicado: (2025)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
por: Tian, Jian, et al.
Publicado: (2025)
por: Tian, Jian, et al.
Publicado: (2025)
Message-Aware Graph Attention Networks for Large-Scale Multi-Robot Path Planning
por: Li, Qingbiao, et al.
Publicado: (2020)
por: Li, Qingbiao, et al.
Publicado: (2020)
Ejemplares similares
-
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
por: Agrawal, Amey, et al.
Publicado: (2026) -
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
por: Lou, Chiheng, et al.
Publicado: (2025) -
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
por: Wu, Bingyang, et al.
Publicado: (2025) -
Locality-aware Fair Scheduling in LLM Serving
por: Cao, Shiyi, et al.
Publicado: (2025) -
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
por: Liu, Yifei, et al.
Publicado: (2025)