Reducing Compute Waste in LLMs through Kernel-Level DVFS
Fuente:
arXiv
Guardado en:
| Autores principales: | Spaan, Jeffrey, Chen, Kuan-Hsun, Varbanescu, Ana-Lucia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
WCDT: Systematic WCET Optimization for Decision Tree Implementations
por: Hölscher, Nils, et al.
Publicado: (2025)
por: Hölscher, Nils, et al.
Publicado: (2025)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
por: Wang, Han, et al.
Publicado: (2026)
por: Wang, Han, et al.
Publicado: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
por: Jaber, Jaber, et al.
Publicado: (2026)
por: Jaber, Jaber, et al.
Publicado: (2026)
Cloud Computing Energy Consumption Prediction Based on Kernel Extreme Learning Machine Algorithm Improved by Vector Weighted Average Algorithm
por: Wang, Yuqing, et al.
Publicado: (2025)
por: Wang, Yuqing, et al.
Publicado: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
por: Ouyang, Anne, et al.
Publicado: (2025)
por: Ouyang, Anne, et al.
Publicado: (2025)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
por: Wang, Liangyu, et al.
Publicado: (2025)
por: Wang, Liangyu, et al.
Publicado: (2025)
A Kernel-Based Approach for Accurate Steady-State Detection in Performance Time Series
por: Beseda, Martin, et al.
Publicado: (2025)
por: Beseda, Martin, et al.
Publicado: (2025)
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
por: Zhang, Yijia, et al.
Publicado: (2024)
por: Zhang, Yijia, et al.
Publicado: (2024)
InTreeger: An End-to-End Framework for Integer-Only Decision Tree Inference
por: Bart, Duncan, et al.
Publicado: (2025)
por: Bart, Duncan, et al.
Publicado: (2025)
SAfEPaTh: A System-Level Approach for Efficient Power and Thermal Estimation of Convolutional Neural Network Accelerator
por: Chen, Yukai, et al.
Publicado: (2024)
por: Chen, Yukai, et al.
Publicado: (2024)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
por: Jiang, Jevin, et al.
Publicado: (2026)
por: Jiang, Jevin, et al.
Publicado: (2026)
Conformer-Based Speech Recognition On Extreme Edge-Computing Devices
por: Xu, Mingbin, et al.
Publicado: (2023)
por: Xu, Mingbin, et al.
Publicado: (2023)
A Structure-Aware Framework for Learning Device Placements on Computation Graphs
por: Duan, Shukai, et al.
Publicado: (2024)
por: Duan, Shukai, et al.
Publicado: (2024)
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
por: Holmes, Connor, et al.
Publicado: (2024)
por: Holmes, Connor, et al.
Publicado: (2024)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
por: Gupta, Ahan, et al.
Publicado: (2023)
por: Gupta, Ahan, et al.
Publicado: (2023)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
por: Dong, Juechu, et al.
Publicado: (2024)
por: Dong, Juechu, et al.
Publicado: (2024)
Towards Computational Performance Engineering for Unsupervised Concept Drift Detection -- Complexities, Benchmarking, Performance Analysis
por: Werner, Elias, et al.
Publicado: (2023)
por: Werner, Elias, et al.
Publicado: (2023)
Accuracy and Consumption analysis from a compressed model by CompactifAI from Multiverse Computing
por: Fovet, Damien, et al.
Publicado: (2025)
por: Fovet, Damien, et al.
Publicado: (2025)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
por: Andrews, Martin, et al.
Publicado: (2025)
por: Andrews, Martin, et al.
Publicado: (2025)
GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
por: Wang, Jiaqi, et al.
Publicado: (2026)
por: Wang, Jiaqi, et al.
Publicado: (2026)
Revisiting Forest Proximities via Sparse Leaf-Incidence Kernels
por: Aumon, Adrien, et al.
Publicado: (2026)
por: Aumon, Adrien, et al.
Publicado: (2026)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
por: Shkolnik, Moran, et al.
Publicado: (2024)
por: Shkolnik, Moran, et al.
Publicado: (2024)
LLMs for Analog Circuit Design Continuum (ACDC)
por: Esfandiari, Yasaman, et al.
Publicado: (2025)
por: Esfandiari, Yasaman, et al.
Publicado: (2025)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
por: Huang, Zixiao, et al.
Publicado: (2025)
por: Huang, Zixiao, et al.
Publicado: (2025)
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
por: Hourri, Younes, et al.
Publicado: (2025)
por: Hourri, Younes, et al.
Publicado: (2025)
Kevin: Multi-Turn RL for Generating CUDA Kernels
por: Baronio, Carlo, et al.
Publicado: (2025)
por: Baronio, Carlo, et al.
Publicado: (2025)
DF-GNN: Dynamic Fusion Framework for Attention Graph Neural Networks on GPUs
por: Liu, Jiahui, et al.
Publicado: (2024)
por: Liu, Jiahui, et al.
Publicado: (2024)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
por: Yun, Vincent-Daniel, et al.
Publicado: (2026)
por: Yun, Vincent-Daniel, et al.
Publicado: (2026)
Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems
por: Panigrahy, Deepak, et al.
Publicado: (2026)
por: Panigrahy, Deepak, et al.
Publicado: (2026)
MoEITS: A Green AI approach for simplifying MoE-LLMs
por: Balderas, Luis, et al.
Publicado: (2026)
por: Balderas, Luis, et al.
Publicado: (2026)
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
por: Mozaffari, Mohammad, et al.
Publicado: (2025)
por: Mozaffari, Mohammad, et al.
Publicado: (2025)
A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
por: Ellis-Mohr, Austin R., et al.
Publicado: (2025)
por: Ellis-Mohr, Austin R., et al.
Publicado: (2025)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
por: Wang, Haoxin, et al.
Publicado: (2025)
por: Wang, Haoxin, et al.
Publicado: (2025)
Anatomizing Deep Learning Inference in Web Browsers
por: Wang, Qipeng, et al.
Publicado: (2024)
por: Wang, Qipeng, et al.
Publicado: (2024)
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
por: Wen, Zhongzhen, et al.
Publicado: (2025)
por: Wen, Zhongzhen, et al.
Publicado: (2025)
Application Research On Real-Time Perception Of Device Performance Status
por: Wang, Zhe, et al.
Publicado: (2024)
por: Wang, Zhe, et al.
Publicado: (2024)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
por: Li, Jianhui, et al.
Publicado: (2023)
por: Li, Jianhui, et al.
Publicado: (2023)
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
por: Liao, Gang, et al.
Publicado: (2025)
por: Liao, Gang, et al.
Publicado: (2025)
DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing
por: Boudaoud, Afif, et al.
Publicado: (2025)
por: Boudaoud, Afif, et al.
Publicado: (2025)
A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations
por: Flavin, Timothy, et al.
Publicado: (2026)
por: Flavin, Timothy, et al.
Publicado: (2026)
Ejemplares similares
-
WCDT: Systematic WCET Optimization for Decision Tree Implementations
por: Hölscher, Nils, et al.
Publicado: (2025) -
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
por: Wang, Han, et al.
Publicado: (2026) -
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
por: Jaber, Jaber, et al.
Publicado: (2026) -
Cloud Computing Energy Consumption Prediction Based on Kernel Extreme Learning Machine Algorithm Improved by Vector Weighted Average Algorithm
por: Wang, Yuqing, et al.
Publicado: (2025) -
KernelBench: Can LLMs Write Efficient GPU Kernels?
por: Ouyang, Anne, et al.
Publicado: (2025)