Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Panigrahy, Deepak, Tyagi, Aakash |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
von: Panigrahy, Deepak, et al.
Veröffentlicht: (2026)
von: Panigrahy, Deepak, et al.
Veröffentlicht: (2026)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
von: Kermani, Arshia, et al.
Veröffentlicht: (2025)
von: Kermani, Arshia, et al.
Veröffentlicht: (2025)
The Power of Training: How Different Neural Network Setups Influence the Energy Demand
von: Geißler, Daniel, et al.
Veröffentlicht: (2024)
von: Geißler, Daniel, et al.
Veröffentlicht: (2024)
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
von: Tran, Nguyen Phuc, et al.
Veröffentlicht: (2025)
von: Tran, Nguyen Phuc, et al.
Veröffentlicht: (2025)
Quantum Neural Networks for Wind Energy Forecasting: A Comparative Study of Performance and Scalability with Classical Models
von: Hangun, Batuhan, et al.
Veröffentlicht: (2025)
von: Hangun, Batuhan, et al.
Veröffentlicht: (2025)
Enhancing Energy-Awareness in Deep Learning through Fine-Grained Energy Measurement
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2023)
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2023)
Throughput Optimization as a Strategic Lever in Large-Scale AI Systems: Evidence from Dataloader and Memory Profiling Innovations
von: Jha, Mayank
Veröffentlicht: (2026)
von: Jha, Mayank
Veröffentlicht: (2026)
On the Sustainability of AI Inferences in the Edge
von: Sobhani, Ghazal, et al.
Veröffentlicht: (2025)
von: Sobhani, Ghazal, et al.
Veröffentlicht: (2025)
The Race to Efficiency: A New Perspective on AI Scaling Laws
von: Lu, Chien-Ping
Veröffentlicht: (2025)
von: Lu, Chien-Ping
Veröffentlicht: (2025)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
MoEITS: A Green AI approach for simplifying MoE-LLMs
von: Balderas, Luis, et al.
Veröffentlicht: (2026)
von: Balderas, Luis, et al.
Veröffentlicht: (2026)
Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
von: Almurshed, Osama, et al.
Veröffentlicht: (2025)
von: Almurshed, Osama, et al.
Veröffentlicht: (2025)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
von: Barad, Haim, et al.
Veröffentlicht: (2023)
von: Barad, Haim, et al.
Veröffentlicht: (2023)
Performance Modeling of Data Storage Systems using Generative Models
von: Al-Maeeni, Abdalaziz Rashid, et al.
Veröffentlicht: (2023)
von: Al-Maeeni, Abdalaziz Rashid, et al.
Veröffentlicht: (2023)
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
von: Liao, Gang, et al.
Veröffentlicht: (2025)
von: Liao, Gang, et al.
Veröffentlicht: (2025)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
von: Yin, Wangsong, et al.
Veröffentlicht: (2025)
von: Yin, Wangsong, et al.
Veröffentlicht: (2025)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
von: Liu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Liu, Jiashuo, et al.
Veröffentlicht: (2025)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
von: Zhao, Yushang, et al.
Veröffentlicht: (2025)
von: Zhao, Yushang, et al.
Veröffentlicht: (2025)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
von: Zheng, Wenhao, et al.
Veröffentlicht: (2025)
von: Zheng, Wenhao, et al.
Veröffentlicht: (2025)
Accelerating AI Performance using Anderson Extrapolation on GPUs
von: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Veröffentlicht: (2024)
von: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Veröffentlicht: (2024)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
von: Jiang, Jevin, et al.
Veröffentlicht: (2026)
von: Jiang, Jevin, et al.
Veröffentlicht: (2026)
Rapid Augmentations for Time Series (RATS): A High-Performance Library for Time Series Augmentation
von: Skaf, Wadie, et al.
Veröffentlicht: (2026)
von: Skaf, Wadie, et al.
Veröffentlicht: (2026)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
von: Knoop, Jonathan, et al.
Veröffentlicht: (2026)
von: Knoop, Jonathan, et al.
Veröffentlicht: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
von: Wu, Wenhao, et al.
Veröffentlicht: (2026)
von: Wu, Wenhao, et al.
Veröffentlicht: (2026)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
von: Shao, Zishan, et al.
Veröffentlicht: (2025)
von: Shao, Zishan, et al.
Veröffentlicht: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2024)
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2024)
Deploying Open-Source Large Language Models: A performance Analysis
von: Bendi-Ouis, Yannis, et al.
Veröffentlicht: (2024)
von: Bendi-Ouis, Yannis, et al.
Veröffentlicht: (2024)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
von: Hossain, Md Arafat, et al.
Veröffentlicht: (2025)
von: Hossain, Md Arafat, et al.
Veröffentlicht: (2025)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
von: Yin, Yishu, et al.
Veröffentlicht: (2025)
von: Yin, Yishu, et al.
Veröffentlicht: (2025)
The Hidden Power of Pure 16-bit Floating-Point Neural Networks
von: Yun, Juyoung, et al.
Veröffentlicht: (2023)
von: Yun, Juyoung, et al.
Veröffentlicht: (2023)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
APOLLO: SGD-like Memory, AdamW-level Performance
von: Zhu, Hanqing, et al.
Veröffentlicht: (2024)
von: Zhu, Hanqing, et al.
Veröffentlicht: (2024)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
von: Qiao, Liang, et al.
Veröffentlicht: (2025)
von: Qiao, Liang, et al.
Veröffentlicht: (2025)
Exchangeability in Neural Network and its Application to Dynamic Pruning
von: Pu, et al.
Veröffentlicht: (2025)
von: Pu, et al.
Veröffentlicht: (2025)
Profiling LoRA/QLoRA Fine-Tuning Efficiency on Consumer GPUs: An RTX 4060 Case Study
von: Avinash, MSR
Veröffentlicht: (2025)
von: Avinash, MSR
Veröffentlicht: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
Machine Learning Methods for Evaluating Public Crisis: Meta-Analysis
von: Okpala, Izunna, et al.
Veröffentlicht: (2023)
von: Okpala, Izunna, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
von: Panigrahy, Deepak, et al.
Veröffentlicht: (2026) -
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
von: Kermani, Arshia, et al.
Veröffentlicht: (2025) -
The Power of Training: How Different Neural Network Setups Influence the Energy Demand
von: Geißler, Daniel, et al.
Veröffentlicht: (2024) -
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
von: Tran, Nguyen Phuc, et al.
Veröffentlicht: (2025) -
Quantum Neural Networks for Wind Energy Forecasting: A Comparative Study of Performance and Scalability with Classical Models
von: Hangun, Batuhan, et al.
Veröffentlicht: (2025)