LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kong, Jie, Wang, Wei, Zhou, Jiehan, Yu, Chen |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
par: Zhang, Yaozheng, et autres
Publié: (2025)
par: Zhang, Yaozheng, et autres
Publié: (2025)
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
par: Kong, Jie, et autres
Publié: (2025)
par: Kong, Jie, et autres
Publié: (2025)
FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
par: Zhang, Yida, et autres
Publié: (2026)
par: Zhang, Yida, et autres
Publié: (2026)
Federated Inference for Heterogeneous LLM Communication and Collaboration
par: Chen, Zihan, et autres
Publié: (2026)
par: Chen, Zihan, et autres
Publié: (2026)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
par: Bambhaniya, Abhimanyu Rajeshkumar, et autres
Publié: (2025)
par: Bambhaniya, Abhimanyu Rajeshkumar, et autres
Publié: (2025)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
par: Xu, Ao, et autres
Publié: (2025)
par: Xu, Ao, et autres
Publié: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
par: Wu, Panlong, et autres
Publié: (2025)
par: Wu, Panlong, et autres
Publié: (2025)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
par: Xu, Yaodan, et autres
Publié: (2025)
par: Xu, Yaodan, et autres
Publié: (2025)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
par: Ye, Fanjiang, et autres
Publié: (2026)
par: Ye, Fanjiang, et autres
Publié: (2026)
KV Cache Compression for Inference Efficiency in LLMs: A Review
par: Liu, Yanyu, et autres
Publié: (2025)
par: Liu, Yanyu, et autres
Publié: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
par: Du, Boxiao, et autres
Publié: (2026)
par: Du, Boxiao, et autres
Publié: (2026)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
par: Chen, Jiu, et autres
Publié: (2026)
par: Chen, Jiu, et autres
Publié: (2026)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
par: Peng, You, et autres
Publié: (2026)
par: Peng, You, et autres
Publié: (2026)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
par: Lu, Yao, et autres
Publié: (2026)
par: Lu, Yao, et autres
Publié: (2026)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
par: Xu, Jiale, et autres
Publié: (2025)
par: Xu, Jiale, et autres
Publié: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
par: Xu, Chuhao, et autres
Publié: (2025)
par: Xu, Chuhao, et autres
Publié: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
par: Wilkins, Grant, et autres
Publié: (2024)
par: Wilkins, Grant, et autres
Publié: (2024)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
par: Yang, Zheming, et autres
Publié: (2025)
par: Yang, Zheming, et autres
Publié: (2025)
Parallax: Efficient LLM Inference Service over Decentralized Environment
par: Tong, Chris, et autres
Publié: (2025)
par: Tong, Chris, et autres
Publié: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
par: He, Yiyuan, et autres
Publié: (2024)
par: He, Yiyuan, et autres
Publié: (2024)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
par: Yu, Shibo, et autres
Publié: (2025)
par: Yu, Shibo, et autres
Publié: (2025)
Efficiently Executing High-throughput Lightweight LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
par: Phung, Thanh Son, et autres
Publié: (2025)
par: Phung, Thanh Son, et autres
Publié: (2025)
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication
par: Rao, Yizhuo, et autres
Publié: (2026)
par: Rao, Yizhuo, et autres
Publié: (2026)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
par: Lee, Sanghyeon, et autres
Publié: (2025)
par: Lee, Sanghyeon, et autres
Publié: (2025)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
par: Guo, Runsheng Benson, et autres
Publié: (2025)
par: Guo, Runsheng Benson, et autres
Publié: (2025)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
par: Jiang, Youhe, et autres
Publié: (2026)
par: Jiang, Youhe, et autres
Publié: (2026)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
par: Da, Wei, et autres
Publié: (2026)
par: Da, Wei, et autres
Publié: (2026)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
par: Liu, Zihan, et autres
Publié: (2025)
par: Liu, Zihan, et autres
Publié: (2025)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
par: Zhang, Songge, et autres
Publié: (2026)
par: Zhang, Songge, et autres
Publié: (2026)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
par: He, Xuan, et autres
Publié: (2025)
par: He, Xuan, et autres
Publié: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
par: Chen, Haoyu, et autres
Publié: (2025)
par: Chen, Haoyu, et autres
Publié: (2025)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
par: Ma, Bin, et autres
Publié: (2026)
par: Ma, Bin, et autres
Publié: (2026)
Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
par: Chen, Xiaopei, et autres
Publié: (2025)
par: Chen, Xiaopei, et autres
Publié: (2025)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
par: Wu, Siyu, et autres
Publié: (2025)
par: Wu, Siyu, et autres
Publié: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
par: He, Wenhao, et autres
Publié: (2026)
par: He, Wenhao, et autres
Publié: (2026)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
par: Liu, Xing, et autres
Publié: (2025)
par: Liu, Xing, et autres
Publié: (2025)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
par: Pang, Bowen, et autres
Publié: (2025)
par: Pang, Bowen, et autres
Publié: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
par: Wang, Chao, et autres
Publié: (2025)
par: Wang, Chao, et autres
Publié: (2025)
Documents similaires
-
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
par: Zhang, Yaozheng, et autres
Publié: (2025) -
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
par: Kong, Jie, et autres
Publié: (2025) -
FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
par: Chen, Huamin, et autres
Publié: (2026) -
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
par: Zhang, Yida, et autres
Publié: (2026) -
Federated Inference for Heterogeneous LLM Communication and Collaboration
par: Chen, Zihan, et autres
Publié: (2026)