LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Da, Wei, Kalyvianaki, Evangelia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
von: Da, Wei, et al.
Veröffentlicht: (2025)
von: Da, Wei, et al.
Veröffentlicht: (2025)
Dodoor: Efficient Randomized Decentralized Scheduling with Load Caching for Heterogeneous Tasks and Clusters
von: Da, Wei, et al.
Veröffentlicht: (2025)
von: Da, Wei, et al.
Veröffentlicht: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
Mitigating context switching in densely packed Linux clusters with Latency-Aware Group Scheduling
von: Isstaif, Al Amjad Tawfiq, et al.
Veröffentlicht: (2025)
von: Isstaif, Al Amjad Tawfiq, et al.
Veröffentlicht: (2025)
An AI-Native Runtime for Multi-Wearable Environments
von: Min, Chulhong, et al.
Veröffentlicht: (2024)
von: Min, Chulhong, et al.
Veröffentlicht: (2024)
Towards an Adaptive Runtime System for Cloud-Native HPC
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
von: Chow, Will
Veröffentlicht: (2025)
von: Chow, Will
Veröffentlicht: (2025)
Fail-Closed Lowering of Resident KV Claims onto LLM Serving Runtimes
von: Stepanek, Lukas
Veröffentlicht: (2026)
von: Stepanek, Lukas
Veröffentlicht: (2026)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
von: Kong, Jie, et al.
Veröffentlicht: (2026)
von: Kong, Jie, et al.
Veröffentlicht: (2026)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
von: Zhan, Huiyou, et al.
Veröffentlicht: (2025)
von: Zhan, Huiyou, et al.
Veröffentlicht: (2025)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
von: Özcan, Miray, et al.
Veröffentlicht: (2025)
von: Özcan, Miray, et al.
Veröffentlicht: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
von: Bournias, Ilias, et al.
Veröffentlicht: (2024)
von: Bournias, Ilias, et al.
Veröffentlicht: (2024)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
Federated Inference for Heterogeneous LLM Communication and Collaboration
von: Chen, Zihan, et al.
Veröffentlicht: (2026)
von: Chen, Zihan, et al.
Veröffentlicht: (2026)
Enabling Dynamic Sparsity in Quantized LLM Inference
von: Wang, Rongxiang, et al.
Veröffentlicht: (2025)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
ACE Runtime - A ZKP-Native Blockchain Runtime with Sub-Second Cryptographic Finality
von: Wang, Jian Sheng
Veröffentlicht: (2026)
von: Wang, Jian Sheng
Veröffentlicht: (2026)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
WANSpec: Leveraging Global Compute Capacity for LLM Inference
von: Martin, Noah, et al.
Veröffentlicht: (2026)
von: Martin, Noah, et al.
Veröffentlicht: (2026)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
von: Wu, Panlong, et al.
Veröffentlicht: (2025)
von: Wu, Panlong, et al.
Veröffentlicht: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
Distributed On-Device LLM Inference With Over-the-Air Computation
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
A Cloud-Native Architecture for Human-in-Control LLM-Assisted OpenSearch in Investigative Settings
von: Puhani, Benjamin, et al.
Veröffentlicht: (2026)
von: Puhani, Benjamin, et al.
Veröffentlicht: (2026)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
von: Da, Wei, et al.
Veröffentlicht: (2025) -
Dodoor: Efficient Randomized Decentralized Scheduling with Load Caching for Heterogeneous Tasks and Clusters
von: Da, Wei, et al.
Veröffentlicht: (2025) -
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025) -
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025) -
Mitigating context switching in densely packed Linux clusters with Latency-Aware Group Scheduling
von: Isstaif, Al Amjad Tawfiq, et al.
Veröffentlicht: (2025)