Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Pang, Bowen, Li, Kai, She, Ruifeng, Wang, Feifan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Llumnix: Dynamic Scheduling for Large Language Model Serving
por: Sun, Biao, et al.
Publicado: (2024)
por: Sun, Biao, et al.
Publicado: (2024)
Serving Large Language Models on Huawei CloudMatrix384
por: Zuo, Pengfei, et al.
Publicado: (2025)
por: Zuo, Pengfei, et al.
Publicado: (2025)
Advancing AI-assisted Hardware Design with Hierarchical Decentralized Training and Personalized Inference-Time Optimization
por: Chen, Hao Mark, et al.
Publicado: (2025)
por: Chen, Hao Mark, et al.
Publicado: (2025)
Online GPU Energy Optimization with Switching-Aware Bandits
por: Xu, Xiongxiao, et al.
Publicado: (2024)
por: Xu, Xiongxiao, et al.
Publicado: (2024)
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
por: Yu, Zhongkai, et al.
Publicado: (2025)
por: Yu, Zhongkai, et al.
Publicado: (2025)
PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing System
por: He, Yintao, et al.
Publicado: (2025)
por: He, Yintao, et al.
Publicado: (2025)
WaferLLM: Large Language Model Inference at Wafer Scale
por: He, Congjie, et al.
Publicado: (2025)
por: He, Congjie, et al.
Publicado: (2025)
Performance Modeling and Workload Analysis of Distributed Large Language Model Training and Inference
por: Kundu, Joyjit, et al.
Publicado: (2024)
por: Kundu, Joyjit, et al.
Publicado: (2024)
ZettaLith: An Architectural Exploration of Extreme-Scale AI Inference Acceleration
por: Silverbrook, Kia
Publicado: (2025)
por: Silverbrook, Kia
Publicado: (2025)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
por: Kakolyris, Andreas Kosmas, et al.
Publicado: (2024)
por: Kakolyris, Andreas Kosmas, et al.
Publicado: (2024)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
por: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Publicado: (2025)
por: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Publicado: (2025)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
por: Bambhaniya, Abhimanyu, et al.
Publicado: (2024)
por: Bambhaniya, Abhimanyu, et al.
Publicado: (2024)
NPU Design for Diffusion Language Model Inference
por: Lou, Binglei, et al.
Publicado: (2026)
por: Lou, Binglei, et al.
Publicado: (2026)
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
por: Odema, Mohanad, et al.
Publicado: (2024)
por: Odema, Mohanad, et al.
Publicado: (2024)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
por: Yeo, Gwangoo, et al.
Publicado: (2024)
por: Yeo, Gwangoo, et al.
Publicado: (2024)
ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving
por: Choi, Yuseon, et al.
Publicado: (2026)
por: Choi, Yuseon, et al.
Publicado: (2026)
PIM-Opt: Demystifying Distributed Optimization Algorithms on a Real-World Processing-In-Memory System
por: Rhyner, Steve, et al.
Publicado: (2024)
por: Rhyner, Steve, et al.
Publicado: (2024)
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
por: Yang, Jinwu, et al.
Publicado: (2026)
por: Yang, Jinwu, et al.
Publicado: (2026)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
por: Liu, Fangxin, et al.
Publicado: (2026)
por: Liu, Fangxin, et al.
Publicado: (2026)
Enabling Accelerators for Graph Computing
por: Shivdikar, Kaustubh
Publicado: (2023)
por: Shivdikar, Kaustubh
Publicado: (2023)
Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
por: Park, Jeongmin Brian, et al.
Publicado: (2023)
por: Park, Jeongmin Brian, et al.
Publicado: (2023)
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
por: Makin, Yashasvi, et al.
Publicado: (2025)
por: Makin, Yashasvi, et al.
Publicado: (2025)
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving
por: Yu, Zhongkai, et al.
Publicado: (2026)
por: Yu, Zhongkai, et al.
Publicado: (2026)
Splitwiser: Efficient LM inference with constrained resources
por: Aali, Asad, et al.
Publicado: (2025)
por: Aali, Asad, et al.
Publicado: (2025)
CUTEv2: Unified and Configurable Matrix Extension for Diverse CPU Architectures with Minimal Design Overhead
por: Ye, Jinpeng, et al.
Publicado: (2026)
por: Ye, Jinpeng, et al.
Publicado: (2026)
Sensitivity-Guided Framework for Pruned and Quantized Reservoir Computing Accelerators
por: Jafari, Atousa, et al.
Publicado: (2026)
por: Jafari, Atousa, et al.
Publicado: (2026)
FlexLink: Boosting your NVLink Bandwidth by 27% without accuracy concern
por: Shen, Ao, et al.
Publicado: (2025)
por: Shen, Ao, et al.
Publicado: (2025)
VLSI Hypergraph Partitioning with Deep Learning
por: Khan, Muhammad Hadir, et al.
Publicado: (2024)
por: Khan, Muhammad Hadir, et al.
Publicado: (2024)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
por: Ding, Jianru, et al.
Publicado: (2026)
por: Ding, Jianru, et al.
Publicado: (2026)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
por: Zhao, Yiren, et al.
Publicado: (2026)
por: Zhao, Yiren, et al.
Publicado: (2026)
Co-design of a novel CMOS highly parallel, low-power, multi-chip neural network accelerator
por: Hokenmaier, W, et al.
Publicado: (2024)
por: Hokenmaier, W, et al.
Publicado: (2024)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
por: Stojkovic, Jovan, et al.
Publicado: (2024)
por: Stojkovic, Jovan, et al.
Publicado: (2024)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems
por: Blanco, Francesco G., et al.
Publicado: (2024)
por: Blanco, Francesco G., et al.
Publicado: (2024)
Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement Learning
por: Russo, Enrico, et al.
Publicado: (2024)
por: Russo, Enrico, et al.
Publicado: (2024)
Efficient, VRAM-Constrained xLM Inference on Clients
por: Ukarande, Aditya, et al.
Publicado: (2026)
por: Ukarande, Aditya, et al.
Publicado: (2026)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
por: Li, Jiaxi, et al.
Publicado: (2025)
por: Li, Jiaxi, et al.
Publicado: (2025)
SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
por: Zhang, Hengrui, et al.
Publicado: (2025)
por: Zhang, Hengrui, et al.
Publicado: (2025)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
por: DeBole, Michael V., et al.
Publicado: (2025)
por: DeBole, Michael V., et al.
Publicado: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
por: Pan, Yudong, et al.
Publicado: (2026)
por: Pan, Yudong, et al.
Publicado: (2026)
Ejemplares similares
-
Llumnix: Dynamic Scheduling for Large Language Model Serving
por: Sun, Biao, et al.
Publicado: (2024) -
Serving Large Language Models on Huawei CloudMatrix384
por: Zuo, Pengfei, et al.
Publicado: (2025) -
Advancing AI-assisted Hardware Design with Hierarchical Decentralized Training and Personalized Inference-Time Optimization
por: Chen, Hao Mark, et al.
Publicado: (2025) -
Online GPU Energy Optimization with Switching-Aware Bandits
por: Xu, Xiongxiao, et al.
Publicado: (2024) -
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
por: Yu, Zhongkai, et al.
Publicado: (2025)