Beyond Test-Time Compute Strategies: Advocating Energy-per-Token in LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Wilhelm, Patrick, Wittkopp, Thorsten, Kao, Odej |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Monitoring Emergent Reward Hacking During Generation via Internal Activations
by: Wilhelm, Patrick, et al.
Published: (2026)
by: Wilhelm, Patrick, et al.
Published: (2026)
LogRCA: Log-based Root Cause Analysis for Distributed Services
by: Wittkopp, Thorsten, et al.
Published: (2024)
by: Wittkopp, Thorsten, et al.
Published: (2024)
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
by: Scheinert, Dominik, et al.
Published: (2026)
by: Scheinert, Dominik, et al.
Published: (2026)
Revisiting Gradient Staleness: Evaluating Distance Metrics for Asynchronous Federated Learning Aggregation
by: Wilhelm, Patrick, et al.
Published: (2026)
by: Wilhelm, Patrick, et al.
Published: (2026)
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
by: Özcan, Miray, et al.
Published: (2025)
by: Özcan, Miray, et al.
Published: (2025)
Comparative Analysis of Large Language Models for the Machine-Assisted Resolution of User Intentions
by: Flerlage, Justus, et al.
Published: (2025)
by: Flerlage, Justus, et al.
Published: (2025)
Carbon-Aware Quality Adaptation for Energy-Intensive Services
by: Wiesner, Philipp, et al.
Published: (2024)
by: Wiesner, Philipp, et al.
Published: (2024)
Noise-aware Client Selection for carbon-efficient Federated Learning via Gradient Norm Thresholding
by: Wilhelm, Patrick, et al.
Published: (2026)
by: Wilhelm, Patrick, et al.
Published: (2026)
Latency and Token-Aware Test-Time Compute
by: Huang, Jenny Y., et al.
Published: (2025)
by: Huang, Jenny Y., et al.
Published: (2025)
Sleep-time Compute: Beyond Inference Scaling at Test-time
by: Lin, Kevin, et al.
Published: (2025)
by: Lin, Kevin, et al.
Published: (2025)
HAMburger: Accelerating LLM Inference via Token Smashing
by: Liu, Jingyu, et al.
Published: (2025)
by: Liu, Jingyu, et al.
Published: (2025)
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
by: Shelat, Shlok, et al.
Published: (2026)
by: Shelat, Shlok, et al.
Published: (2026)
Extending Token Computation for LLM Reasoning
by: Liao, Bingli, et al.
Published: (2024)
by: Liao, Bingli, et al.
Published: (2024)
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
Empirical Evidences for the Effects of Feature Diversity in Open Set Recognition and Continual Learning
by: Xu, Jiawen, et al.
Published: (2025)
by: Xu, Jiawen, et al.
Published: (2025)
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
Do Large Language Models Advocate for Inferentialism?
by: Arai, Yuzuki, et al.
Published: (2024)
by: Arai, Yuzuki, et al.
Published: (2024)
Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies
by: Wang, Junlin, et al.
Published: (2024)
by: Wang, Junlin, et al.
Published: (2024)
SDSAT: Accelerating LLM Inference through Speculative Decoding with Semantic Adaptive Tokens
by: Liu, Chengbo, et al.
Published: (2024)
by: Liu, Chengbo, et al.
Published: (2024)
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Softmax Attention with Constant Cost per Token
by: Heinsen, Franz A.
Published: (2024)
by: Heinsen, Franz A.
Published: (2024)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
by: Yang, Wenkai, et al.
Published: (2025)
by: Yang, Wenkai, et al.
Published: (2025)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
by: Bi, Zhenni, et al.
Published: (2024)
by: Bi, Zhenni, et al.
Published: (2024)
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
by: Dong, Ximing, et al.
Published: (2026)
by: Dong, Ximing, et al.
Published: (2026)
A layered architecture for log analysis in complex IT systems
by: Wittkopp, Thorsten
Published: (2025)
by: Wittkopp, Thorsten
Published: (2025)
Beyond Literal Token Overlap: Token Alignability for Multilinguality
by: Hämmerl, Katharina, et al.
Published: (2025)
by: Hämmerl, Katharina, et al.
Published: (2025)
SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
by: Long, Lingkun, et al.
Published: (2025)
by: Long, Lingkun, et al.
Published: (2025)
Mask Tokens as Prophet: Fine-Grained Cache Eviction for Efficient dLLM Inference
by: Huang, Jianuo, et al.
Published: (2025)
by: Huang, Jianuo, et al.
Published: (2025)
LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
by: Ding, Yiran, et al.
Published: (2024)
by: Ding, Yiran, et al.
Published: (2024)
Expanding Computation Spaces of LLMs at Inference Time
by: Jang, Yoonna, et al.
Published: (2025)
by: Jang, Yoonna, et al.
Published: (2025)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
by: Han, Chao, et al.
Published: (2025)
by: Han, Chao, et al.
Published: (2025)
Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
by: Qin, Zongyue, et al.
Published: (2024)
by: Qin, Zongyue, et al.
Published: (2024)
Rethinking Agentic Workflows: Evaluating Inference-Based Test-Time Scaling Strategies in Text2SQL Tasks
by: Guo, Jiajing, et al.
Published: (2025)
by: Guo, Jiajing, et al.
Published: (2025)
A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning
by: Ji, Yixin, et al.
Published: (2025)
by: Ji, Yixin, et al.
Published: (2025)
Moving Beyond Marginal Carbon Intensity: A Poor Metric for Both Carbon Accounting and Grid Flexibility
by: Wiesner, Philipp, et al.
Published: (2025)
by: Wiesner, Philipp, et al.
Published: (2025)
TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
by: Koike-Akino, Toshiaki, et al.
Published: (2026)
by: Koike-Akino, Toshiaki, et al.
Published: (2026)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024)
by: Fu, Qichen, et al.
Published: (2024)
Compute Optimal Tokenization
by: Limisiewicz, Tomasz, et al.
Published: (2026)
by: Limisiewicz, Tomasz, et al.
Published: (2026)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
A Survey on LLM Inference-Time Self-Improvement
by: Dong, Xiangjue, et al.
Published: (2024)
by: Dong, Xiangjue, et al.
Published: (2024)
Similar Items
-
Monitoring Emergent Reward Hacking During Generation via Internal Activations
by: Wilhelm, Patrick, et al.
Published: (2026) -
LogRCA: Log-based Root Cause Analysis for Distributed Services
by: Wittkopp, Thorsten, et al.
Published: (2024) -
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
by: Scheinert, Dominik, et al.
Published: (2026) -
Revisiting Gradient Staleness: Evaluating Distance Metrics for Asynchronous Federated Learning Aggregation
by: Wilhelm, Patrick, et al.
Published: (2026) -
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
by: Özcan, Miray, et al.
Published: (2025)