Position: LLM Inference Should Be Evaluated as Energy-to-Token Production
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiang, Yuan, Shimiao, Tang, Zhenheng, Dong, Peijie, Zhao, Kaiyong, Wang, Qiang, Li, Bo, Chu, Xiaowen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
by: Tang, Zhenheng, et al.
Published: (2025)
by: Tang, Zhenheng, et al.
Published: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
by: Tang, Zhenheng, et al.
Published: (2024)
by: Tang, Zhenheng, et al.
Published: (2024)
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025)
by: Wu, Panlong, et al.
Published: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
by: Wang, Yuxin, et al.
Published: (2024)
by: Wang, Yuxin, et al.
Published: (2024)
FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion
by: Tang, Zhenheng, et al.
Published: (2024)
by: Tang, Zhenheng, et al.
Published: (2024)
Bandwidth-Aware and Overlap-Weighted Compression for Communication-Efficient Federated Learning
by: Tang, Zichen, et al.
Published: (2024)
by: Tang, Zichen, et al.
Published: (2024)
A Cloud-based Real-time Probabilistic Remaining Useful Life (RUL) Estimation using the Sequential Monte Carlo (SMC) Method
by: Lyathakula, Karthik Reddy, et al.
Published: (2024)
by: Lyathakula, Karthik Reddy, et al.
Published: (2024)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
by: Wilkins, Grant, et al.
Published: (2024)
by: Wilkins, Grant, et al.
Published: (2024)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
by: Khoshsirat, Aria, et al.
Published: (2024)
by: Khoshsirat, Aria, et al.
Published: (2024)
FedImpro: Measuring and Improving Client Update in Federated Learning
by: Tang, Zhenheng, et al.
Published: (2024)
by: Tang, Zhenheng, et al.
Published: (2024)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
by: Zhang, Yuning, et al.
Published: (2025)
by: Zhang, Yuning, et al.
Published: (2025)
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
by: Özcan, Miray, et al.
Published: (2025)
by: Özcan, Miray, et al.
Published: (2025)
A GPU-boosted high-performance multi-working condition joint analysis framework for predicting dynamics of textured axial piston pump
by: Yao, Xin, et al.
Published: (2025)
by: Yao, Xin, et al.
Published: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
by: Wang, Yuxin, et al.
Published: (2023)
by: Wang, Yuxin, et al.
Published: (2023)
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
by: Lin, Wenxiang, et al.
Published: (2026)
by: Lin, Wenxiang, et al.
Published: (2026)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
by: Xu, Chuhao, et al.
Published: (2025)
by: Xu, Chuhao, et al.
Published: (2025)
Characterizing GPU Energy Usage in Exascale-Ready Portable Science Applications
by: Godoy, William F., et al.
Published: (2025)
by: Godoy, William F., et al.
Published: (2025)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
by: Kang, Xueze, et al.
Published: (2025)
by: Kang, Xueze, et al.
Published: (2025)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
by: Lai, Ruiqi, et al.
Published: (2025)
by: Lai, Ruiqi, et al.
Published: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
by: Sun, Mingyu, et al.
Published: (2025)
by: Sun, Mingyu, et al.
Published: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
by: Bian, Zhuohang, et al.
Published: (2026)
by: Bian, Zhuohang, et al.
Published: (2026)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
by: Dong, Jiangwen, et al.
Published: (2025)
by: Dong, Jiangwen, et al.
Published: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
by: Wu, Yu, et al.
Published: (2025)
by: Wu, Yu, et al.
Published: (2025)
Parallax: Efficient LLM Inference Service over Decentralized Environment
by: Tong, Chris, et al.
Published: (2025)
by: Tong, Chris, et al.
Published: (2025)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
by: Chen, Jiu, et al.
Published: (2026)
by: Chen, Jiu, et al.
Published: (2026)
Reduced and mixed precision turbulent flow simulations using explicit finite difference schemes
by: Siklósi, Bálint, et al.
Published: (2025)
by: Siklósi, Bálint, et al.
Published: (2025)
In-Memory Non-Binary LDPC Decoding
by: Ferraz, Oscar, et al.
Published: (2025)
by: Ferraz, Oscar, et al.
Published: (2025)
A parallel implementation of reduced-order modeling of large-scale systems
by: Farcas, Ionut-Gabriel, et al.
Published: (2025)
by: Farcas, Ionut-Gabriel, et al.
Published: (2025)
Reclaiming Idle CPU Cycles on Kubernetes: Sparse-Domain Multiplexing for Concurrent MPI-CFD Simulations
by: Xie, Tianfang
Published: (2026)
by: Xie, Tianfang
Published: (2026)
Making Tax Smart: Feasibility of Distributed Ledger Technology for building tax compliance functionality to Central Bank Digital Currency
by: Louvieris, Panos, et al.
Published: (2024)
by: Louvieris, Panos, et al.
Published: (2024)
GPU-accelerated Linear Algebra for Coupled Solvers in Industrial CFD Applications with OpenFOAM
by: Oliani, Stefano, et al.
Published: (2024)
by: Oliani, Stefano, et al.
Published: (2024)
Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels
by: Yang, Shaoliang, et al.
Published: (2026)
by: Yang, Shaoliang, et al.
Published: (2026)
Optimizing the Weather Research and Forecasting Model with OpenMP Offload and Codee
by: Chayanon, et al.
Published: (2024)
by: Chayanon, et al.
Published: (2024)
Mass Matrix Assembly on Tensor Cores for Implicit Particle-In-Cell Methods
by: Pennati, Luca, et al.
Published: (2026)
by: Pennati, Luca, et al.
Published: (2026)
A GPU-based Compressible Combustion Solver for Applications Exhibiting Disparate Space and Time Scales
by: Carreon, Anthony, et al.
Published: (2025)
by: Carreon, Anthony, et al.
Published: (2025)
Level set-based inverse homogenisation of three-dimensional piezoelectric materials
by: Wegert, Zachary J., et al.
Published: (2024)
by: Wegert, Zachary J., et al.
Published: (2024)
Intertemporal Pricing of Time-Bound Stablecoins: Measuring and Controlling the Liquidity-of-Time Premium
by: Borjigin, Ailiya, et al.
Published: (2025)
by: Borjigin, Ailiya, et al.
Published: (2025)
A Portable Multi-GPU Solver for Collisional Plasmas with Coulombic Interactions
by: Almgren-Bell, James, et al.
Published: (2025)
by: Almgren-Bell, James, et al.
Published: (2025)
Smoothed aggregation algebraic multigrid for problems with heterogeneous and anisotropic materials
by: Firmbach, Max, et al.
Published: (2026)
by: Firmbach, Max, et al.
Published: (2026)
Similar Items
-
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
by: Tang, Zhenheng, et al.
Published: (2025) -
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025) -
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
by: Tang, Zhenheng, et al.
Published: (2024) -
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025) -
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
by: Wang, Yuxin, et al.
Published: (2024)