Input-Dependent Power Usage in GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Gregersen, Theo, Patel, Pratyush, Choukse, Esha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
Natural Language Query to Configuration for Retrieval Agents
by: Pan, Melissa Z., et al.
Published: (2026)
by: Pan, Melissa Z., et al.
Published: (2026)
Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency
by: Hong, Colin, et al.
Published: (2025)
by: Hong, Colin, et al.
Published: (2025)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
Towards Resource-Efficient Compound AI Systems
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
by: Qiu, Haoran, et al.
Published: (2026)
by: Qiu, Haoran, et al.
Published: (2026)
Synthesizing Access Control Policies using Large Language Models
by: Vatsa, Adarsh, et al.
Published: (2025)
by: Vatsa, Adarsh, et al.
Published: (2025)
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
by: Priyanshu, Aman, et al.
Published: (2026)
by: Priyanshu, Aman, et al.
Published: (2026)
Splitwise: Efficient generative LLM inference using phase splitting
by: Patel, Pratyush, et al.
Published: (2023)
by: Patel, Pratyush, et al.
Published: (2023)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
by: Qiu, Haoran, et al.
Published: (2025)
by: Qiu, Haoran, et al.
Published: (2025)
CuRLA: Curriculum Learning Based Deep Reinforcement Learning for Autonomous Driving
by: Uppuluri, Bhargava, et al.
Published: (2025)
by: Uppuluri, Bhargava, et al.
Published: (2025)
Grokking as a Variance-Limited Phase Transition: Spectral Gating and the Epsilon-Stability Threshold
by: Acharya, Pratyush, et al.
Published: (2026)
by: Acharya, Pratyush, et al.
Published: (2026)
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
by: Liu, Yuhan, et al.
Published: (2024)
by: Liu, Yuhan, et al.
Published: (2024)
LTL learning on GPUs
by: Valizadeh, Mojtaba, et al.
Published: (2024)
by: Valizadeh, Mojtaba, et al.
Published: (2024)
Fully-fused Multi-Layer Perceptrons on Intel Data Center GPUs
by: Yuan, Kai, et al.
Published: (2024)
by: Yuan, Kai, et al.
Published: (2024)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
by: Cui, Shengkun, et al.
Published: (2025)
by: Cui, Shengkun, et al.
Published: (2025)
Asynchronous Tool Usage for Real-Time Agents
by: Ginart, Antonio A., et al.
Published: (2024)
by: Ginart, Antonio A., et al.
Published: (2024)
Mind the Gap: Comparing Model- vs Agentic-Level Red Teaming with Action-Graph Observability on GPT-OSS-20B
by: Wicaksono, Ilham, et al.
Published: (2025)
by: Wicaksono, Ilham, et al.
Published: (2025)
TempOpt -- Unsupervised Alarm Relation Learning for Telecommunication Networks
by: Sampath, Sathiyanaryanan, et al.
Published: (2025)
by: Sampath, Sathiyanaryanan, et al.
Published: (2025)
Concepts Whisper While Syntax Shouts: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations
by: Acharya, Pratyush, et al.
Published: (2026)
by: Acharya, Pratyush, et al.
Published: (2026)
Electromagnetic Simulations of Antennas on GPUs for Machine Learning Applications
by: Temiz, Murat, et al.
Published: (2025)
by: Temiz, Murat, et al.
Published: (2025)
Syntactic Evolution in Language Usage
by: Kumar, Surbhit
Published: (2025)
by: Kumar, Surbhit
Published: (2025)
Usage Governance Advisor: From Intent to AI Governance
by: Daly, Elizabeth M., et al.
Published: (2024)
by: Daly, Elizabeth M., et al.
Published: (2024)
RoundTable: Leveraging Dynamic Schema and Contextual Autocomplete for Enhanced Query Precision in Tabular Question Answering
by: Kumar, Pratyush, et al.
Published: (2024)
by: Kumar, Pratyush, et al.
Published: (2024)
Holistic Surgical Phase Recognition with Hierarchical Input Dependent State Space Models
by: Wu, Haoyang, et al.
Published: (2025)
by: Wu, Haoyang, et al.
Published: (2025)
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
by: Zhong, Zhengjia, et al.
Published: (2026)
by: Zhong, Zhengjia, et al.
Published: (2026)
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
by: Yu, Seungmin, et al.
Published: (2024)
by: Yu, Seungmin, et al.
Published: (2024)
LiNR: Model Based Neural Retrieval on GPUs at LinkedIn
by: Borisyuk, Fedor, et al.
Published: (2024)
by: Borisyuk, Fedor, et al.
Published: (2024)
How Much Progress Has There Been in NVIDIA Datacenter GPUs?
by: Del Sozzo, Emanuele, et al.
Published: (2026)
by: Del Sozzo, Emanuele, et al.
Published: (2026)
Memorization Sinks: Isolating Memorization during LLM Training
by: Ghosal, Gaurav R., et al.
Published: (2025)
by: Ghosal, Gaurav R., et al.
Published: (2025)
Hard Examples Are All You Need: Maximizing GRPO Post-Training Under Annotation Budgets
by: Pikus, Benjamin, et al.
Published: (2025)
by: Pikus, Benjamin, et al.
Published: (2025)
Truck Parking Usage Prediction with Decomposed Graph Neural Networks
by: Tamaru, Rei, et al.
Published: (2024)
by: Tamaru, Rei, et al.
Published: (2024)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
Understanding Usage and Engagement in AI-Powered Scientific Research Tools: The Asta Interaction Dataset
by: Haddad, Dany, et al.
Published: (2026)
by: Haddad, Dany, et al.
Published: (2026)
Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
by: Liu, Yi
Published: (2026)
by: Liu, Yi
Published: (2026)
Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers
by: Pratyush, Spandan
Published: (2026)
by: Pratyush, Spandan
Published: (2026)
DynamicGate MLP Conditional Computation via Learned Structural Dropout and Input Dependent Gating for Functional Plasticity
by: Choi, Yong Il
Published: (2026)
by: Choi, Yong Il
Published: (2026)
Similar Items
-
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
by: Stojkovic, Jovan, et al.
Published: (2025) -
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024) -
Natural Language Query to Configuration for Retrieval Agents
by: Pan, Melissa Z., et al.
Published: (2026) -
Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency
by: Hong, Colin, et al.
Published: (2025) -
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
by: Stojkovic, Jovan, et al.
Published: (2024)