Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
Fuente:
arXiv
Saved in:
| Main Authors: | Nicolae, Radu, Iosup, Alexandru, Trivedi, Animesh, Donkervliet, Jesse |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
OpenDT: Exploring Datacenter Performance and Sustainability with a Self-Calibrating Digital Twin
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
by: Niewenhuis, Dante, et al.
Published: (2026)
by: Niewenhuis, Dante, et al.
Published: (2026)
Toward Sustainability-Aware LLM Inference on Edge Clusters
by: Rajashekar, Kolichala, et al.
Published: (2025)
by: Rajashekar, Kolichala, et al.
Published: (2025)
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
by: Suman, Shekhar, et al.
Published: (2026)
by: Suman, Shekhar, et al.
Published: (2026)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
by: Chu, Xiaoyu, et al.
Published: (2025)
by: Chu, Xiaoyu, et al.
Published: (2025)
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
by: Battaglini-Fischer, Sándor, et al.
Published: (2025)
by: Battaglini-Fischer, Sándor, et al.
Published: (2025)
KV Cache Compression for Inference Efficiency in LLMs: A Review
by: Liu, Yanyu, et al.
Published: (2025)
by: Liu, Yanyu, et al.
Published: (2025)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
by: Lee, Sanghyeon, et al.
Published: (2025)
by: Lee, Sanghyeon, et al.
Published: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
by: Arif, Moiz, et al.
Published: (2026)
by: Arif, Moiz, et al.
Published: (2026)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
by: Tummalapalli, Pranay, et al.
Published: (2026)
by: Tummalapalli, Pranay, et al.
Published: (2026)
PARSIR: a Package for Effective Parallel Discrete Event Simulation on Multi-processor Machines
by: Quaglia, Francesco
Published: (2024)
by: Quaglia, Francesco
Published: (2024)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
by: Zhang, Yuning, et al.
Published: (2025)
by: Zhang, Yuning, et al.
Published: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
by: Karfakis, George, et al.
Published: (2025)
by: Karfakis, George, et al.
Published: (2025)
Efficient Data-Parallel Continual Learning with Asynchronous Distributed Rehearsal Buffers
by: Bouvier, Thomas, et al.
Published: (2024)
by: Bouvier, Thomas, et al.
Published: (2024)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
by: Wang, Weiye, et al.
Published: (2026)
by: Wang, Weiye, et al.
Published: (2026)
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025)
by: Wu, Panlong, et al.
Published: (2025)
AIReSim: A Discrete Event Simulator for Large-scale AI Cluster Reliability Modeling
by: Pattabiraman, Karthik, et al.
Published: (2026)
by: Pattabiraman, Karthik, et al.
Published: (2026)
Harnessing Your DRAM and SSD for Sustainable and Accessible LLM Inference with Mixed-Precision and Multi-level Caching
by: Peng, Jie, et al.
Published: (2024)
by: Peng, Jie, et al.
Published: (2024)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
by: Nian, Sean, et al.
Published: (2026)
by: Nian, Sean, et al.
Published: (2026)
KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
by: Arya, Mayank, et al.
Published: (2025)
by: Arya, Mayank, et al.
Published: (2025)
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
by: Afroz, Sabiha, et al.
Published: (2025)
by: Afroz, Sabiha, et al.
Published: (2025)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
by: Lin, Shouxu, et al.
Published: (2026)
by: Lin, Shouxu, et al.
Published: (2026)
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
by: Özcan, Miray, et al.
Published: (2025)
by: Özcan, Miray, et al.
Published: (2025)
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
by: Gossman, Mikaila J., et al.
Published: (2025)
by: Gossman, Mikaila J., et al.
Published: (2025)
SCAREY: Location-Aware Service Lifecycle Management
by: Horvath, Kurt, et al.
Published: (2025)
by: Horvath, Kurt, et al.
Published: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
by: Wang, Qipeng
Published: (2026)
by: Wang, Qipeng
Published: (2026)
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
by: Ruan, Chaoyi, et al.
Published: (2025)
by: Ruan, Chaoyi, et al.
Published: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
by: Fang, Shaoke, et al.
Published: (2026)
by: Fang, Shaoke, et al.
Published: (2026)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
by: Jiang, Chaoyi, et al.
Published: (2024)
by: Jiang, Chaoyi, et al.
Published: (2024)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
by: Kim, Joon Ha, et al.
Published: (2026)
by: Kim, Joon Ha, et al.
Published: (2026)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
Experimental Analysis of Server-Side Caching for Web Performance
by: Umar, Mohammad, et al.
Published: (2026)
by: Umar, Mohammad, et al.
Published: (2026)
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
by: Xia, Tian, et al.
Published: (2025)
by: Xia, Tian, et al.
Published: (2025)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
by: Argerich, Mauricio Fadel, et al.
Published: (2026)
by: Argerich, Mauricio Fadel, et al.
Published: (2026)
Similar Items
-
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
by: Nicolae, Radu, et al.
Published: (2026) -
OpenDT: Exploring Datacenter Performance and Sustainability with a Self-Calibrating Digital Twin
by: Nicolae, Radu, et al.
Published: (2026) -
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
by: Niewenhuis, Dante, et al.
Published: (2026) -
Toward Sustainability-Aware LLM Inference on Edge Clusters
by: Rajashekar, Kolichala, et al.
Published: (2025) -
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
by: Suman, Shekhar, et al.
Published: (2026)