TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Feiyang, Bian, Zhuohang, Duan, Guoyang, Xu, Tianle, Wu, Junchi, Ma, Teng, Yao, Yongqiang, Gong, Ruihao, Zhuo, Youwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
by: Bian, Zhuohang, et al.
Published: (2026)
by: Bian, Zhuohang, et al.
Published: (2026)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
by: Bian, Zhuohang, et al.
Published: (2025)
by: Bian, Zhuohang, et al.
Published: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025)
by: Wu, Panlong, et al.
Published: (2025)
Toward Systems Foundations for Agentic Exploration
by: Xu, Jiakai, et al.
Published: (2025)
by: Xu, Jiakai, et al.
Published: (2025)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
by: Dai, Yuntao, et al.
Published: (2026)
by: Dai, Yuntao, et al.
Published: (2026)
Bridging Simulation and Silicon: A Study of RISC-V Hardware and FireSim Simulation
by: Barai, Atanu, et al.
Published: (2025)
by: Barai, Atanu, et al.
Published: (2025)
Benchmarking Compound AI Applications for Hardware-Software Co-Design
by: Samuthrsindh, Paramuth, et al.
Published: (2026)
by: Samuthrsindh, Paramuth, et al.
Published: (2026)
Past-Future Scheduler for LLM Serving under SLA Guarantees
by: Gong, Ruihao, et al.
Published: (2025)
by: Gong, Ruihao, et al.
Published: (2025)
PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads
by: Wu, Xin, et al.
Published: (2026)
by: Wu, Xin, et al.
Published: (2026)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
by: Wu, Jingfeng, et al.
Published: (2024)
by: Wu, Jingfeng, et al.
Published: (2024)
Twinning for Space-Air-Ground-Sea Integrated Networks: Beyond Conventional Digital Twin Towards Goal-Oriented Semantic Twin
by: Qiu, Yifei, et al.
Published: (2025)
by: Qiu, Yifei, et al.
Published: (2025)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
by: Wang, Hansheng, et al.
Published: (2024)
by: Wang, Hansheng, et al.
Published: (2024)
Enabling Dynamic Sparsity in Quantized LLM Inference
by: Wang, Rongxiang, et al.
Published: (2025)
by: Wang, Rongxiang, et al.
Published: (2025)
SimDC: A High-Fidelity Device Simulation Platform for Device-Cloud Collaborative Computing
by: Pei, Ruiguang, et al.
Published: (2025)
by: Pei, Ruiguang, et al.
Published: (2025)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
by: Wu, Tian, et al.
Published: (2025)
by: Wu, Tian, et al.
Published: (2025)
An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
by: Hakim, Sheikh Azizul, et al.
Published: (2025)
by: Hakim, Sheikh Azizul, et al.
Published: (2025)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
by: Liang, Yan, et al.
Published: (2026)
by: Liang, Yan, et al.
Published: (2026)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
by: Yoo, Jinsun, et al.
Published: (2026)
by: Yoo, Jinsun, et al.
Published: (2026)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
by: Li, Haley, et al.
Published: (2026)
by: Li, Haley, et al.
Published: (2026)
Age-Aware Partial Gradient Update Strategy for Federated Learning Over the Air
by: Du, Ruihao, et al.
Published: (2025)
by: Du, Ruihao, et al.
Published: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
by: Yu, Dianhai, et al.
Published: (2022)
by: Yu, Dianhai, et al.
Published: (2022)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
by: Chen, Liangkun, et al.
Published: (2025)
by: Chen, Liangkun, et al.
Published: (2025)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
by: Wang, Yingping, et al.
Published: (2026)
by: Wang, Yingping, et al.
Published: (2026)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
by: Wu, Yu, et al.
Published: (2025)
by: Wu, Yu, et al.
Published: (2025)
On-the-fly Communication-and-Computing to Enable Representation Learning for Distributed Point Clouds
by: Chen, Xu, et al.
Published: (2024)
by: Chen, Xu, et al.
Published: (2024)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
by: Lee, Sunjung, et al.
Published: (2026)
by: Lee, Sunjung, et al.
Published: (2026)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
by: Yu, Minchen, et al.
Published: (2023)
by: Yu, Minchen, et al.
Published: (2023)
DRackSim: Simulator for Rack-scale Memory Disaggregation
by: Puri, Amit, et al.
Published: (2023)
by: Puri, Amit, et al.
Published: (2023)
Cloud Native System for LLM Inference Serving
by: Xu, Minxian, et al.
Published: (2025)
by: Xu, Minxian, et al.
Published: (2025)
LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure
by: Cho, Jaehong, et al.
Published: (2025)
by: Cho, Jaehong, et al.
Published: (2025)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
by: Lin, Shouxu, et al.
Published: (2026)
by: Lin, Shouxu, et al.
Published: (2026)
Towards Efficient Verification of Parallel Applications with Mc SimGrid
by: Laurent, Matthieu, et al.
Published: (2025)
by: Laurent, Matthieu, et al.
Published: (2025)
Hierarchical Observe-Orient-Decide-Act Enabled UAV Swarms in Uncertain Environments: Frameworks, Potentials, and Challenges
by: Jia, Ziye, et al.
Published: (2026)
by: Jia, Ziye, et al.
Published: (2026)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
by: Qu, Huanyu, et al.
Published: (2025)
by: Qu, Huanyu, et al.
Published: (2025)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
by: Zhang, Songge, et al.
Published: (2026)
by: Zhang, Songge, et al.
Published: (2026)
ESTA: An Efficient Spatial-Temporal Range Aggregation Query Processing Algorithm for UAV Networks
by: Liu, Liang, et al.
Published: (2023)
by: Liu, Liang, et al.
Published: (2023)
Federated Semi-Supervised and Semi-Asynchronous Learning for Anomaly Detection in IoT Networks
by: Zhai, Wenbin, et al.
Published: (2023)
by: Zhai, Wenbin, et al.
Published: (2023)
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
by: Hanafy, Walid A., et al.
Published: (2025)
by: Hanafy, Walid A., et al.
Published: (2025)
Similar Items
-
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
by: Bian, Zhuohang, et al.
Published: (2026) -
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
by: Bian, Zhuohang, et al.
Published: (2025) -
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025) -
Toward Systems Foundations for Agentic Exploration
by: Xu, Jiakai, et al.
Published: (2025) -
Accelerating OpenPangu Inference on NPU via Speculative Decoding
by: Dai, Yuntao, et al.
Published: (2026)