Elastic On-Device LLM Service
Fuente:
arXiv
Salvato in:
| Autori principali: | Yin, Wangsong, Yi, Rongjie, Xu, Daliang, Huang, Gang, Xu, Mengwei, Liu, Xuanzhe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Survey of Resource-efficient LLM and Multimodal Foundation Models
di: Xu, Mengwei, et al.
Pubblicazione: (2024)
di: Xu, Mengwei, et al.
Pubblicazione: (2024)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
di: Wang, Shiju, et al.
Pubblicazione: (2025)
di: Wang, Shiju, et al.
Pubblicazione: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
di: Ye, Chenhao, et al.
Pubblicazione: (2026)
di: Ye, Chenhao, et al.
Pubblicazione: (2026)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
di: Liu, Xing, et al.
Pubblicazione: (2025)
di: Liu, Xing, et al.
Pubblicazione: (2025)
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
di: Wang, Li, et al.
Pubblicazione: (2024)
di: Wang, Li, et al.
Pubblicazione: (2024)
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
di: Lv, Jiaqi, et al.
Pubblicazione: (2025)
di: Lv, Jiaqi, et al.
Pubblicazione: (2025)
Deploying Foundation Model Powered Agent Services: A Survey
di: Xu, Wenchao, et al.
Pubblicazione: (2024)
di: Xu, Wenchao, et al.
Pubblicazione: (2024)
xLLM Technical Report
di: Liu, Tongxuan, et al.
Pubblicazione: (2025)
di: Liu, Tongxuan, et al.
Pubblicazione: (2025)
Design a Win-Win Strategy That Is Fair to Both Service Providers and Tasks When Rejection Is Not an Option
di: Trabelsi, Yohai, et al.
Pubblicazione: (2024)
di: Trabelsi, Yohai, et al.
Pubblicazione: (2024)
High-Throughput LLM inference on Heterogeneous Clusters
di: Xiong, Yi, et al.
Pubblicazione: (2025)
di: Xiong, Yi, et al.
Pubblicazione: (2025)
Balanced and Elastic End-to-end Training of Dynamic LLMs
di: Wahib, Mohamed, et al.
Pubblicazione: (2025)
di: Wahib, Mohamed, et al.
Pubblicazione: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
di: Zhao, Yihao, et al.
Pubblicazione: (2025)
di: Zhao, Yihao, et al.
Pubblicazione: (2025)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
di: Liu, Man, et al.
Pubblicazione: (2026)
di: Liu, Man, et al.
Pubblicazione: (2026)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
di: Wang, Haodong, et al.
Pubblicazione: (2025)
di: Wang, Haodong, et al.
Pubblicazione: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
di: Li, Rongzhi, et al.
Pubblicazione: (2025)
di: Li, Rongzhi, et al.
Pubblicazione: (2025)
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
di: Xi, Shaoke, et al.
Pubblicazione: (2026)
di: Xi, Shaoke, et al.
Pubblicazione: (2026)
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
di: Yang, Xinjun, et al.
Pubblicazione: (2025)
di: Yang, Xinjun, et al.
Pubblicazione: (2025)
Dynamic Resource Allocation for Virtual Machine Migration Optimization using Machine Learning
di: Gong, Yulu, et al.
Pubblicazione: (2024)
di: Gong, Yulu, et al.
Pubblicazione: (2024)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
di: Spoczynski, Marcin, et al.
Pubblicazione: (2026)
di: Spoczynski, Marcin, et al.
Pubblicazione: (2026)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
di: Huang, Tao, et al.
Pubblicazione: (2024)
di: Huang, Tao, et al.
Pubblicazione: (2024)
Revisiting Parameter Server in LLM Post-Training
di: Wan, Xinyi, et al.
Pubblicazione: (2026)
di: Wan, Xinyi, et al.
Pubblicazione: (2026)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
di: Chen, Huamin, et al.
Pubblicazione: (2026)
di: Chen, Huamin, et al.
Pubblicazione: (2026)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
Robust Synchronisation for Federated Learning in The Face of Correlated Device Failure
di: Behfar, Stefan, et al.
Pubblicazione: (2026)
di: Behfar, Stefan, et al.
Pubblicazione: (2026)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
di: Xu, Jiale, et al.
Pubblicazione: (2025)
di: Xu, Jiale, et al.
Pubblicazione: (2025)
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
di: Hao, Zixu, et al.
Pubblicazione: (2025)
di: Hao, Zixu, et al.
Pubblicazione: (2025)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
di: Lyu, Hongtao, et al.
Pubblicazione: (2025)
di: Lyu, Hongtao, et al.
Pubblicazione: (2025)
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
di: Wang, Tuowei, et al.
Pubblicazione: (2025)
di: Wang, Tuowei, et al.
Pubblicazione: (2025)
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
di: Luo, Xinhao, et al.
Pubblicazione: (2025)
di: Luo, Xinhao, et al.
Pubblicazione: (2025)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
di: Wu, Bingyang, et al.
Pubblicazione: (2025)
di: Wu, Bingyang, et al.
Pubblicazione: (2025)
Synera: Synergistic LLM Serving across Device and Cloud at Scale
di: Wang, Genglin, et al.
Pubblicazione: (2025)
di: Wang, Genglin, et al.
Pubblicazione: (2025)
PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding
di: McDanel, Bradley, et al.
Pubblicazione: (2025)
di: McDanel, Bradley, et al.
Pubblicazione: (2025)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
A Nonlinear Hash-based Optimization Method for SpMV on GPUs
di: Yan, Chen, et al.
Pubblicazione: (2025)
di: Yan, Chen, et al.
Pubblicazione: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
di: Wu, Linyu, et al.
Pubblicazione: (2025)
di: Wu, Linyu, et al.
Pubblicazione: (2025)
Simplifying Root Cause Analysis in Kubernetes with StateGraph and LLM
di: Xiang, Yong, et al.
Pubblicazione: (2025)
di: Xiang, Yong, et al.
Pubblicazione: (2025)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained Devices
di: Chen, Fahao, et al.
Pubblicazione: (2025)
di: Chen, Fahao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Survey of Resource-efficient LLM and Multimodal Foundation Models
di: Xu, Mengwei, et al.
Pubblicazione: (2024) -
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
di: Wang, Shiju, et al.
Pubblicazione: (2025) -
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
di: Li, Xiangyu, et al.
Pubblicazione: (2025) -
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
di: Ye, Chenhao, et al.
Pubblicazione: (2026) -
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
di: Liu, Xing, et al.
Pubblicazione: (2025)