Gespeichert in:
| Hauptverfasser: | Chandrasekar, Ashok, Kramberger, Jason |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.24217 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025)
von: Park, Jay H., et al.
Veröffentlicht: (2025)
AI Benchmarks and Datasets for LLM Evaluation
von: Ivanov, Todor, et al.
Veröffentlicht: (2024)
von: Ivanov, Todor, et al.
Veröffentlicht: (2024)
LAPS: A Length-Aware-Prefill LLM Serving System
von: She, Jianshu, et al.
Veröffentlicht: (2026)
von: She, Jianshu, et al.
Veröffentlicht: (2026)
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
von: He, Zifan, et al.
Veröffentlicht: (2026)
von: He, Zifan, et al.
Veröffentlicht: (2026)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
Seesaw: High-throughput LLM Inference via Model Re-sharding
von: Su, Qidong, et al.
Veröffentlicht: (2025)
von: Su, Qidong, et al.
Veröffentlicht: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
On Evaluating Performance of LLM Inference Serving Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
von: Li, Wanqian, et al.
Veröffentlicht: (2026)
von: Li, Wanqian, et al.
Veröffentlicht: (2026)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks
von: Nader, Noujoud, et al.
Veröffentlicht: (2025)
von: Nader, Noujoud, et al.
Veröffentlicht: (2025)
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
von: Hsieh, Chia-chi, et al.
Veröffentlicht: (2026)
von: Hsieh, Chia-chi, et al.
Veröffentlicht: (2026)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
von: Cho, Jaehong, et al.
Veröffentlicht: (2024)
von: Cho, Jaehong, et al.
Veröffentlicht: (2024)
Frontier: Simulating the Next Generation of LLM Inference Systems
von: Feng, Yicheng, et al.
Veröffentlicht: (2025)
von: Feng, Yicheng, et al.
Veröffentlicht: (2025)
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
von: Luo, Xinhao, et al.
Veröffentlicht: (2025)
von: Luo, Xinhao, et al.
Veröffentlicht: (2025)
Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
Benchmarking of CPU-intensive Stream Data Processing in The Edge Computing Systems
von: Szydlo, Tomasz, et al.
Veröffentlicht: (2025)
von: Szydlo, Tomasz, et al.
Veröffentlicht: (2025)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
von: Nagaitsev, Kirill, et al.
Veröffentlicht: (2025)
von: Nagaitsev, Kirill, et al.
Veröffentlicht: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
von: Yu, Dianhai, et al.
Veröffentlicht: (2022)
von: Yu, Dianhai, et al.
Veröffentlicht: (2022)
GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
von: VG, Jithin, et al.
Veröffentlicht: (2025)
von: VG, Jithin, et al.
Veröffentlicht: (2025)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware
von: Khalil, Alex, et al.
Veröffentlicht: (2025)
von: Khalil, Alex, et al.
Veröffentlicht: (2025)
TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
von: Niu, Chenxu, et al.
Veröffentlicht: (2025)
von: Niu, Chenxu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026) -
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025) -
AI Benchmarks and Datasets for LLM Evaluation
von: Ivanov, Todor, et al.
Veröffentlicht: (2024) -
LAPS: A Length-Aware-Prefill LLM Serving System
von: She, Jianshu, et al.
Veröffentlicht: (2026) -
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
von: He, Zifan, et al.
Veröffentlicht: (2026)