GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Niam, Arefin, Kosar, Tevfik, Nine, M. S. Q. Zulkar
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913061920571392
author Niam, Arefin
Kosar, Tevfik
Nine, M. S. Q. Zulkar
author_facet Niam, Arefin
Kosar, Tevfik
Nine, M. S. Q. Zulkar
contents Distributed GNN training is dominated by remote feature fetching, which can be very costly. Multi-hop neighborhood sampling crosses partition boundaries and triggers fine-grained RPCs whose fixed initiation cost and GPU-stall latency waste energy. Prior systems try to reduce this overhead with presampling and static caching, but cache policies cannot react to runtime network variation. We show that under time-varying congestion, static caching can increase energy by up to 45% because a fixed rebuild schedule is insufficient. We present GreenDyGNN, which formulates cache window management as a sequential decision problem. GreenDyGNN performs intra-epoch cache rebuilds and uses a Double-DQN agent, trained in a calibrated simulator with domain-randomized congestion, to adapt rebuild window size and per-owner cache allocation at each boundary. An asynchronous double-buffered pipeline makes adaptation effectively free. Under congestion, GreenDyGNN cuts total energy by up to 43% over Default DGL and 4-24% over the best static policy, while closely matching the optimum under clean conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23139
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
Niam, Arefin
Kosar, Tevfik
Nine, M. S. Q. Zulkar
Distributed, Parallel, and Cluster Computing
Distributed GNN training is dominated by remote feature fetching, which can be very costly. Multi-hop neighborhood sampling crosses partition boundaries and triggers fine-grained RPCs whose fixed initiation cost and GPU-stall latency waste energy. Prior systems try to reduce this overhead with presampling and static caching, but cache policies cannot react to runtime network variation. We show that under time-varying congestion, static caching can increase energy by up to 45% because a fixed rebuild schedule is insufficient. We present GreenDyGNN, which formulates cache window management as a sequential decision problem. GreenDyGNN performs intra-epoch cache rebuilds and uses a Double-DQN agent, trained in a calibrated simulator with domain-randomized congestion, to adapt rebuild window size and per-owner cache allocation at each boundary. An asynchronous double-buffered pipeline makes adaptation effectively free. Under congestion, GreenDyGNN cuts total energy by up to 43% over Default DGL and 4-24% over the best static policy, while closely matching the optimum under clean conditions.
title GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2604.23139