Frontier: Simulating the Next Generation of LLM Inference Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Feng, Yicheng, Tan, Xin, Sew, Kin Hang, Jiang, Yimin, Zhu, Yibo, Xu, Hong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
di: Feng, Yicheng, et al.
Pubblicazione: (2026)
di: Feng, Yicheng, et al.
Pubblicazione: (2026)
Teola: Towards End-to-End Optimization of LLM-based Applications
di: Tan, Xin, et al.
Pubblicazione: (2024)
di: Tan, Xin, et al.
Pubblicazione: (2024)
Accelerating Distributed MoE Training and Inference with Lina
di: Li, Jiamin, et al.
Pubblicazione: (2022)
di: Li, Jiamin, et al.
Pubblicazione: (2022)
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
di: Tan, Xin, et al.
Pubblicazione: (2026)
di: Tan, Xin, et al.
Pubblicazione: (2026)
Seesaw: High-throughput LLM Inference via Model Re-sharding
di: Su, Qidong, et al.
Pubblicazione: (2025)
di: Su, Qidong, et al.
Pubblicazione: (2025)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
di: Xue, Zhenliang, et al.
Pubblicazione: (2025)
di: Xue, Zhenliang, et al.
Pubblicazione: (2025)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
di: Kim, Joon Ha, et al.
Pubblicazione: (2026)
di: Kim, Joon Ha, et al.
Pubblicazione: (2026)
LLM Inference Serving: Survey of Recent Advances and Opportunities
di: Li, Baolin, et al.
Pubblicazione: (2024)
di: Li, Baolin, et al.
Pubblicazione: (2024)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
di: Miyashita, Yusuke, et al.
Pubblicazione: (2024)
di: Miyashita, Yusuke, et al.
Pubblicazione: (2024)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
di: Cho, Jaehong, et al.
Pubblicazione: (2024)
di: Cho, Jaehong, et al.
Pubblicazione: (2024)
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
di: Chandrasekar, Ashok, et al.
Pubblicazione: (2026)
di: Chandrasekar, Ashok, et al.
Pubblicazione: (2026)
Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
di: Wang, Zhanghan, et al.
Pubblicazione: (2025)
di: Wang, Zhanghan, et al.
Pubblicazione: (2025)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
di: Chen, Huamin, et al.
Pubblicazione: (2026)
di: Chen, Huamin, et al.
Pubblicazione: (2026)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
di: Liu, Xing, et al.
Pubblicazione: (2025)
di: Liu, Xing, et al.
Pubblicazione: (2025)
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
di: The AIBrix Team, et al.
Pubblicazione: (2025)
di: The AIBrix Team, et al.
Pubblicazione: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
Accelerating LLM Inference with Precomputed Query Storage
di: Park, Jay H., et al.
Pubblicazione: (2025)
di: Park, Jay H., et al.
Pubblicazione: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
di: Ramachandran, Arun, et al.
Pubblicazione: (2025)
di: Ramachandran, Arun, et al.
Pubblicazione: (2025)
Decentralized AI: Permissionless LLM Inference on POKT Network
di: Olshansky, Daniel, et al.
Pubblicazione: (2024)
di: Olshansky, Daniel, et al.
Pubblicazione: (2024)
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
di: Xu, Lang, et al.
Pubblicazione: (2025)
di: Xu, Lang, et al.
Pubblicazione: (2025)
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
di: Luo, Xinhao, et al.
Pubblicazione: (2025)
di: Luo, Xinhao, et al.
Pubblicazione: (2025)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
di: Lyu, Hongtao, et al.
Pubblicazione: (2025)
di: Lyu, Hongtao, et al.
Pubblicazione: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
di: Li, Rongzhi, et al.
Pubblicazione: (2025)
di: Li, Rongzhi, et al.
Pubblicazione: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
di: Wu, Linyu, et al.
Pubblicazione: (2025)
di: Wu, Linyu, et al.
Pubblicazione: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
di: Yu, Dianhai, et al.
Pubblicazione: (2022)
di: Yu, Dianhai, et al.
Pubblicazione: (2022)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
di: Chen, Le, et al.
Pubblicazione: (2025)
di: Chen, Le, et al.
Pubblicazione: (2025)
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
di: Li, Yufei, et al.
Pubblicazione: (2025)
di: Li, Yufei, et al.
Pubblicazione: (2025)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
di: Song, Jingwei, et al.
Pubblicazione: (2025)
di: Song, Jingwei, et al.
Pubblicazione: (2025)
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
di: Argerich, Mauricio Fadel, et al.
Pubblicazione: (2026)
di: Argerich, Mauricio Fadel, et al.
Pubblicazione: (2026)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
di: Li, Wanqian, et al.
Pubblicazione: (2026)
di: Li, Wanqian, et al.
Pubblicazione: (2026)
On Evaluating Performance of LLM Inference Serving Systems
di: Agrawal, Amey, et al.
Pubblicazione: (2025)
di: Agrawal, Amey, et al.
Pubblicazione: (2025)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
Elastic On-Device LLM Service
di: Yin, Wangsong, et al.
Pubblicazione: (2024)
di: Yin, Wangsong, et al.
Pubblicazione: (2024)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
di: Cho, Jaehong, et al.
Pubblicazione: (2026)
di: Cho, Jaehong, et al.
Pubblicazione: (2026)
Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems
di: Alshaer, Samer, et al.
Pubblicazione: (2025)
di: Alshaer, Samer, et al.
Pubblicazione: (2025)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
di: Zhang, Ping, et al.
Pubblicazione: (2024)
di: Zhang, Ping, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
di: Feng, Yicheng, et al.
Pubblicazione: (2026) -
Teola: Towards End-to-End Optimization of LLM-based Applications
di: Tan, Xin, et al.
Pubblicazione: (2024) -
Accelerating Distributed MoE Training and Inference with Lina
di: Li, Jiamin, et al.
Pubblicazione: (2022) -
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
di: Tan, Xin, et al.
Pubblicazione: (2026) -
Seesaw: High-throughput LLM Inference via Model Re-sharding
di: Su, Qidong, et al.
Pubblicazione: (2025)