A Unified Performance-Cost-Efficiency Model for Large-Scale AI Training on GPU and TPU Accelerators

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Revista, Zen, IA, 10
Format: Recurso digital
Veröffentlicht: Zenodo 2025
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901646178516992
author Revista, Zen
IA, 10
author_facet Revista, Zen
IA, 10
contents The escalating demand for large-scale Artificial Intelligence (AI) models has propelled a critical need for efficient and cost-effective training methodologies. Modern AI training, particularly for deep learning models like large language models and vision transformers, relies heavily on specialized hardware accelerators such as Graphics Processing Units (GPUs) and Tensor Processing Units (TPUs). While these accelerators offer significant computational power, optimizing their utilization involves complex trade-offs between raw performance, monetary cost, and energy efficiency. Current approaches often optimize for one metric in isolation, leading to suboptimal resource allocation and increased operational expenditure. This paper introduces a novel, unified performance-cost-efficiency model designed to systematically evaluate and predict the integrated impact of hardware choice, model architecture, dataset size, and training parameters on these three crucial metrics. The model synthesizes architectural characteristics of leading GPUs and TPUs with real-world cloud pricing structures and power consumption profiles. Through a rigorous methodological framework encompassing analytical modeling and empirical validation across diverse AI workloads, our research demonstrates the model's capability to accurately forecast optimal accelerator configurations. The findings provide actionable insights for researchers and practitioners, enabling informed decisions that balance computational speed with economic viability and environmental responsibility, thereby facilitating more sustainable and efficient large-scale AI development.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17827602
institution Zenodo
language
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle A Unified Performance-Cost-Efficiency Model for Large-Scale AI Training on GPU and TPU Accelerators
Revista, Zen
IA, 10
The escalating demand for large-scale Artificial Intelligence (AI) models has propelled a critical need for efficient and cost-effective training methodologies. Modern AI training, particularly for deep learning models like large language models and vision transformers, relies heavily on specialized hardware accelerators such as Graphics Processing Units (GPUs) and Tensor Processing Units (TPUs). While these accelerators offer significant computational power, optimizing their utilization involves complex trade-offs between raw performance, monetary cost, and energy efficiency. Current approaches often optimize for one metric in isolation, leading to suboptimal resource allocation and increased operational expenditure. This paper introduces a novel, unified performance-cost-efficiency model designed to systematically evaluate and predict the integrated impact of hardware choice, model architecture, dataset size, and training parameters on these three crucial metrics. The model synthesizes architectural characteristics of leading GPUs and TPUs with real-world cloud pricing structures and power consumption profiles. Through a rigorous methodological framework encompassing analytical modeling and empirical validation across diverse AI workloads, our research demonstrates the model's capability to accurately forecast optimal accelerator configurations. The findings provide actionable insights for researchers and practitioners, enabling informed decisions that balance computational speed with economic viability and environmental responsibility, thereby facilitating more sustainable and efficient large-scale AI development.
title A Unified Performance-Cost-Efficiency Model for Large-Scale AI Training on GPU and TPU Accelerators
url https://doi.org/10.5281/zenodo.17827602