A Unified Performance-Cost-Efficiency Model for Large-Scale AI Training on GPU and TPU Accelerators
Fuente:
Zenodo
Gespeichert in:
| Hauptverfasser: | , |
|---|---|
| Format: | Recurso digital |
| Veröffentlicht: |
Zenodo
2025
|
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901646178516992 |
|---|---|
| author | Revista, Zen IA, 10 |
| author_facet | Revista, Zen IA, 10 |
| contents | The escalating demand for large-scale Artificial Intelligence (AI) models has propelled a critical need for efficient and cost-effective training methodologies. Modern AI training, particularly for deep learning models like large language models and vision transformers, relies heavily on specialized hardware accelerators such as Graphics Processing Units (GPUs) and Tensor Processing Units (TPUs). While these accelerators offer significant computational power, optimizing their utilization involves complex trade-offs between raw performance, monetary cost, and energy efficiency. Current approaches often optimize for one metric in isolation, leading to suboptimal resource allocation and increased operational expenditure. This paper introduces a novel, unified performance-cost-efficiency model designed to systematically evaluate and predict the integrated impact of hardware choice, model architecture, dataset size, and training parameters on these three crucial metrics. The model synthesizes architectural characteristics of leading GPUs and TPUs with real-world cloud pricing structures and power consumption profiles. Through a rigorous methodological framework encompassing analytical modeling and empirical validation across diverse AI workloads, our research demonstrates the model's capability to accurately forecast optimal accelerator configurations. The findings provide actionable insights for researchers and practitioners, enabling informed decisions that balance computational speed with economic viability and environmental responsibility, thereby facilitating more sustainable and efficient large-scale AI development. |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17827602 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | A Unified Performance-Cost-Efficiency Model for Large-Scale AI Training on GPU and TPU Accelerators Revista, Zen IA, 10 The escalating demand for large-scale Artificial Intelligence (AI) models has propelled a critical need for efficient and cost-effective training methodologies. Modern AI training, particularly for deep learning models like large language models and vision transformers, relies heavily on specialized hardware accelerators such as Graphics Processing Units (GPUs) and Tensor Processing Units (TPUs). While these accelerators offer significant computational power, optimizing their utilization involves complex trade-offs between raw performance, monetary cost, and energy efficiency. Current approaches often optimize for one metric in isolation, leading to suboptimal resource allocation and increased operational expenditure. This paper introduces a novel, unified performance-cost-efficiency model designed to systematically evaluate and predict the integrated impact of hardware choice, model architecture, dataset size, and training parameters on these three crucial metrics. The model synthesizes architectural characteristics of leading GPUs and TPUs with real-world cloud pricing structures and power consumption profiles. Through a rigorous methodological framework encompassing analytical modeling and empirical validation across diverse AI workloads, our research demonstrates the model's capability to accurately forecast optimal accelerator configurations. The findings provide actionable insights for researchers and practitioners, enabling informed decisions that balance computational speed with economic viability and environmental responsibility, thereby facilitating more sustainable and efficient large-scale AI development. |
| title | A Unified Performance-Cost-Efficiency Model for Large-Scale AI Training on GPU and TPU Accelerators |
| url | https://doi.org/10.5281/zenodo.17827602 |