TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Niu, Chenxu, Zhang, Wei, Li, Jie, Zhao, Yongjian, Wang, Tongyang, Wang, Xi, Chen, Yong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Argus: Token Aware Distributed LLM Inference Optimization
di: Wu, Panlong, et al.
Pubblicazione: (2025)
di: Wu, Panlong, et al.
Pubblicazione: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
di: Arya, Mayank, et al.
Pubblicazione: (2025)
di: Arya, Mayank, et al.
Pubblicazione: (2025)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
di: Chen, Huamin, et al.
Pubblicazione: (2026)
di: Chen, Huamin, et al.
Pubblicazione: (2026)
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
di: Kong, Jie, et al.
Pubblicazione: (2026)
di: Kong, Jie, et al.
Pubblicazione: (2026)
Element and Everything Tokens: Two-Tier Architecture for Mobilizing Alternative Assets
di: Borjigin, Ailiya, et al.
Pubblicazione: (2025)
di: Borjigin, Ailiya, et al.
Pubblicazione: (2025)
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
di: Özcan, Miray, et al.
Pubblicazione: (2025)
di: Özcan, Miray, et al.
Pubblicazione: (2025)
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
di: Yuan, Yichao, et al.
Pubblicazione: (2026)
di: Yuan, Yichao, et al.
Pubblicazione: (2026)
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
di: Wilkins, Grant, et al.
Pubblicazione: (2026)
di: Wilkins, Grant, et al.
Pubblicazione: (2026)
Power Aware Dynamic Reallocation For Inference
di: Jiang, Yiwei, et al.
Pubblicazione: (2026)
di: Jiang, Yiwei, et al.
Pubblicazione: (2026)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
di: Hankendi, Can, et al.
Pubblicazione: (2026)
di: Hankendi, Can, et al.
Pubblicazione: (2026)
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
di: Chandrasekar, Ashok, et al.
Pubblicazione: (2026)
di: Chandrasekar, Ashok, et al.
Pubblicazione: (2026)
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
di: Liu, Guangyi, et al.
Pubblicazione: (2026)
di: Liu, Guangyi, et al.
Pubblicazione: (2026)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
di: Proaño, Andrès Rubio, et al.
Pubblicazione: (2024)
di: Proaño, Andrès Rubio, et al.
Pubblicazione: (2024)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
di: Spoczynski, Marcin, et al.
Pubblicazione: (2026)
di: Spoczynski, Marcin, et al.
Pubblicazione: (2026)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
di: Argerich, Mauricio Fadel, et al.
Pubblicazione: (2026)
di: Argerich, Mauricio Fadel, et al.
Pubblicazione: (2026)
Exploring User Acceptance of Blockchain-Based Student Certificate Sharing System: A Study on Non Fungible Token (NFT) Utilization
di: Khati, Prakhyat, et al.
Pubblicazione: (2024)
di: Khati, Prakhyat, et al.
Pubblicazione: (2024)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
di: Zhao, Yiren, et al.
Pubblicazione: (2026)
di: Zhao, Yiren, et al.
Pubblicazione: (2026)
Seesaw: High-throughput LLM Inference via Model Re-sharding
di: Su, Qidong, et al.
Pubblicazione: (2025)
di: Su, Qidong, et al.
Pubblicazione: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
di: Chen, Haoyu, et al.
Pubblicazione: (2025)
di: Chen, Haoyu, et al.
Pubblicazione: (2025)
Deploying Foundation Model Powered Agent Services: A Survey
di: Xu, Wenchao, et al.
Pubblicazione: (2024)
di: Xu, Wenchao, et al.
Pubblicazione: (2024)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
di: Abramovich, Talor, et al.
Pubblicazione: (2026)
di: Abramovich, Talor, et al.
Pubblicazione: (2026)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
di: Song, Jingwei, et al.
Pubblicazione: (2025)
di: Song, Jingwei, et al.
Pubblicazione: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
EfiMon: A Process Analyser for Granular Power Consumption Prediction
di: León-Vega, Luis G., et al.
Pubblicazione: (2024)
di: León-Vega, Luis G., et al.
Pubblicazione: (2024)
CheckMate: LLM-Powered Approximate Intermittent Computing
di: Sayyid-Ali, Abdur-Rahman Ibrahim, et al.
Pubblicazione: (2024)
di: Sayyid-Ali, Abdur-Rahman Ibrahim, et al.
Pubblicazione: (2024)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
di: Lyu, Hongtao, et al.
Pubblicazione: (2025)
di: Lyu, Hongtao, et al.
Pubblicazione: (2025)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
di: Zhang, Yida, et al.
Pubblicazione: (2026)
di: Zhang, Yida, et al.
Pubblicazione: (2026)
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
di: Bhatia, Nidhi, et al.
Pubblicazione: (2025)
di: Bhatia, Nidhi, et al.
Pubblicazione: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
di: Li, Rongzhi, et al.
Pubblicazione: (2025)
di: Li, Rongzhi, et al.
Pubblicazione: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
di: Wang, Weiye, et al.
Pubblicazione: (2026)
di: Wang, Weiye, et al.
Pubblicazione: (2026)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
di: Zhao, Alan, et al.
Pubblicazione: (2026)
di: Zhao, Alan, et al.
Pubblicazione: (2026)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
di: Xu, Guanbin, et al.
Pubblicazione: (2026)
di: Xu, Guanbin, et al.
Pubblicazione: (2026)
Accelerating LLM Inference with Precomputed Query Storage
di: Park, Jay H., et al.
Pubblicazione: (2025)
di: Park, Jay H., et al.
Pubblicazione: (2025)
AI Benchmarks and Datasets for LLM Evaluation
di: Ivanov, Todor, et al.
Pubblicazione: (2024)
di: Ivanov, Todor, et al.
Pubblicazione: (2024)
GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
di: VG, Jithin, et al.
Pubblicazione: (2025)
di: VG, Jithin, et al.
Pubblicazione: (2025)
Enabling Dynamic Sparsity in Quantized LLM Inference
di: Wang, Rongxiang, et al.
Pubblicazione: (2025)
di: Wang, Rongxiang, et al.
Pubblicazione: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Argus: Token Aware Distributed LLM Inference Optimization
di: Wu, Panlong, et al.
Pubblicazione: (2025) -
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025) -
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
di: Arya, Mayank, et al.
Pubblicazione: (2025) -
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
di: Chen, Huamin, et al.
Pubblicazione: (2026) -
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
di: Wilkins, Grant, et al.
Pubblicazione: (2024)