Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yicheng, Tan, Xin, Deng, Yangtao, Jiang, Yimin, Zhu, Yibo, Xu, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Frontier: Simulating the Next Generation of LLM Inference Systems
by: Feng, Yicheng, et al.
Published: (2025)
by: Feng, Yicheng, et al.
Published: (2025)
Teola: Towards End-to-End Optimization of LLM-based Applications
by: Tan, Xin, et al.
Published: (2024)
by: Tan, Xin, et al.
Published: (2024)
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
by: Tan, Xin, et al.
Published: (2026)
by: Tan, Xin, et al.
Published: (2026)
Accelerating Distributed MoE Training and Inference with Lina
by: Li, Jiamin, et al.
Published: (2022)
by: Li, Jiamin, et al.
Published: (2022)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
by: Chen, Le, et al.
Published: (2025)
by: Chen, Le, et al.
Published: (2025)
Seesaw: High-throughput LLM Inference via Model Re-sharding
by: Su, Qidong, et al.
Published: (2025)
by: Su, Qidong, et al.
Published: (2025)
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
by: The AIBrix Team, et al.
Published: (2025)
by: The AIBrix Team, et al.
Published: (2025)
Lodestar: An Online-Learning LLM Inference Router
by: Lim, Gangmuk, et al.
Published: (2026)
by: Lim, Gangmuk, et al.
Published: (2026)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
by: Xue, Zhenliang, et al.
Published: (2025)
by: Xue, Zhenliang, et al.
Published: (2025)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
by: Deng, Xiumei, et al.
Published: (2025)
by: Deng, Xiumei, et al.
Published: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
by: Jiang, Xuanlin, et al.
Published: (2024)
by: Jiang, Xuanlin, et al.
Published: (2024)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
by: Kim, Joon Ha, et al.
Published: (2026)
by: Kim, Joon Ha, et al.
Published: (2026)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
by: Gond, Raja, et al.
Published: (2026)
by: Gond, Raja, et al.
Published: (2026)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
by: Wu, Linyu, et al.
Published: (2025)
by: Wu, Linyu, et al.
Published: (2025)
LLM Inference Serving: Survey of Recent Advances and Opportunities
by: Li, Baolin, et al.
Published: (2024)
by: Li, Baolin, et al.
Published: (2024)
Niyama : Breaking the Silos of LLM Inference Serving
by: Goel, Kanishk, et al.
Published: (2025)
by: Goel, Kanishk, et al.
Published: (2025)
On Evaluating Performance of LLM Inference Serving Systems
by: Agrawal, Amey, et al.
Published: (2025)
by: Agrawal, Amey, et al.
Published: (2025)
Calibre: Towards Fair and Accurate Personalized Federated Learning with Self-Supervised Learning
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
by: Cho, Jaehong, et al.
Published: (2024)
by: Cho, Jaehong, et al.
Published: (2024)
Towards the Next Frontier of LLMs, Training on Private Data: A Cross-Domain Benchmark for Federated Fine-Tuning
by: Jimenez-Gutierrez, Daniel M., et al.
Published: (2026)
by: Jimenez-Gutierrez, Daniel M., et al.
Published: (2026)
MatKV: Trading Compute for Flash Storage in LLM Inference
by: Shin, Kun-Woo, et al.
Published: (2025)
by: Shin, Kun-Woo, et al.
Published: (2025)
DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
by: Peng, Xian, et al.
Published: (2025)
by: Peng, Xian, et al.
Published: (2025)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
by: Chen, Huamin, et al.
Published: (2026)
by: Chen, Huamin, et al.
Published: (2026)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
by: Lu, Yunchi, et al.
Published: (2025)
by: Lu, Yunchi, et al.
Published: (2025)
RelayGR: Scaling Long-Sequence Generative Recommendation via Cross-Stage Relay-Race Inference
by: Wang, Jiarui, et al.
Published: (2026)
by: Wang, Jiarui, et al.
Published: (2026)
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
by: Yu, Donglin
Published: (2026)
by: Yu, Donglin
Published: (2026)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
by: Deshmukh, Dhruv, et al.
Published: (2025)
by: Deshmukh, Dhruv, et al.
Published: (2025)
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
by: Deng, Yangtao, et al.
Published: (2025)
by: Deng, Yangtao, et al.
Published: (2025)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
by: Yu, Jiahuan, et al.
Published: (2026)
by: Yu, Jiahuan, et al.
Published: (2026)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
by: Rhee, Myunghyun, et al.
Published: (2025)
by: Rhee, Myunghyun, et al.
Published: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
by: Behera, Adarsh Prasad, et al.
Published: (2025)
by: Behera, Adarsh Prasad, et al.
Published: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
by: Zheng, Wanyi, et al.
Published: (2025)
by: Zheng, Wanyi, et al.
Published: (2025)
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
by: Wang, Tuowei, et al.
Published: (2025)
by: Wang, Tuowei, et al.
Published: (2025)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
by: Levine, Reese, et al.
Published: (2026)
by: Levine, Reese, et al.
Published: (2026)
Accelerating LLM Inference with Precomputed Query Storage
by: Park, Jay H., et al.
Published: (2025)
by: Park, Jay H., et al.
Published: (2025)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
by: Hu, Huanqi, et al.
Published: (2025)
by: Hu, Huanqi, et al.
Published: (2025)
Distributed Graph Neural Network Inference With Just-In-Time Compilation For Industry-Scale Graphs
by: Wu, Xiabao, et al.
Published: (2025)
by: Wu, Xiabao, et al.
Published: (2025)
Similar Items
-
Frontier: Simulating the Next Generation of LLM Inference Systems
by: Feng, Yicheng, et al.
Published: (2025) -
Teola: Towards End-to-End Optimization of LLM-based Applications
by: Tan, Xin, et al.
Published: (2024) -
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
by: Tan, Xin, et al.
Published: (2026) -
Accelerating Distributed MoE Training and Inference with Lina
by: Li, Jiamin, et al.
Published: (2022) -
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
by: Chen, Le, et al.
Published: (2025)