Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Ruofan, Chung, Jae-Won, Chowdhury, Mosharaf
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908786289016832
author Wu, Ruofan
Chung, Jae-Won
Chowdhury, Mosharaf
author_facet Wu, Ruofan
Chung, Jae-Won
Chowdhury, Mosharaf
contents The computing demand of AI is growing at an unprecedented rate, but energy supply is not keeping pace. As a result, energy has become an expensive, contended resource that requires explicit management and optimization. Although recent works have made significant progress in large model training optimization, they focus only on a single aspect of energy consumption: dynamic or static energy. We find that fine-grained kernel scheduling and frequency scaling jointly and interdependently impact both dynamic and static energy consumption. Based on this finding, we design Kareus, a training system that pushes the time--energy tradeoff frontier by optimizing both aspects. Kareus decomposes the intractable joint optimization problem into local, partition-based subproblems. It then uses a multi-pass multi-objective optimization algorithm to find execution schedules that push the time--energy tradeoff frontier. Compared to the state of the art, Kareus reduces training energy by up to 28.3% at the same training time, or reduces training time by up to 27.5% at the same energy consumption.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17654
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training
Wu, Ruofan
Chung, Jae-Won
Chowdhury, Mosharaf
Machine Learning
Distributed, Parallel, and Cluster Computing
The computing demand of AI is growing at an unprecedented rate, but energy supply is not keeping pace. As a result, energy has become an expensive, contended resource that requires explicit management and optimization. Although recent works have made significant progress in large model training optimization, they focus only on a single aspect of energy consumption: dynamic or static energy. We find that fine-grained kernel scheduling and frequency scaling jointly and interdependently impact both dynamic and static energy consumption. Based on this finding, we design Kareus, a training system that pushes the time--energy tradeoff frontier by optimizing both aspects. Kareus decomposes the intractable joint optimization problem into local, partition-based subproblems. It then uses a multi-pass multi-objective optimization algorithm to find execution schedules that push the time--energy tradeoff frontier. Compared to the state of the art, Kareus reduces training energy by up to 28.3% at the same training time, or reduces training time by up to 27.5% at the same energy consumption.
title Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2601.17654