Efficient allocation of image recognition and LLM tasks on multi-GPU system

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lawenda, Marcin, Samborski, Krzesimir, Khloponin, Kyrylo, Szustak, Łukasz
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910883859398656
author Lawenda, Marcin
Samborski, Krzesimir
Khloponin, Kyrylo
Szustak, Łukasz
author_facet Lawenda, Marcin
Samborski, Krzesimir
Khloponin, Kyrylo
Szustak, Łukasz
contents This work is concerned with the evaluation of the performance of parallelization of learning and tuning processes for image classification and large language models. For machine learning model in image recognition, various parallelization methods are developed based on different hardware and software scenarios: simple data parallelism, distributed data parallelism, and distributed processing. A detailed description of presented strategies is given, highlighting the challenges and benefits of their application. Furthermore, the impact of different dataset types on the tuning process of large language models is investigated. Experiments show to what extent the task type affects the iteration time in a multi-GPU environment, offering valuable insights into the optimal data utilization strategies to improve model performance. Furthermore, this study leverages the built-in parallelization mechanisms of PyTorch that can facilitate these tasks. Furthermore, performance profiling is incorporated into the study to thoroughly evaluate the impact of memory and communication operations during the training/tuning procedure. Test scenarios are developed and tested with numerous benchmarks on the NVIDIA H100 architecture showing efficiency through selected metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15252
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient allocation of image recognition and LLM tasks on multi-GPU system
Lawenda, Marcin
Samborski, Krzesimir
Khloponin, Kyrylo
Szustak, Łukasz
Distributed, Parallel, and Cluster Computing
Performance
This work is concerned with the evaluation of the performance of parallelization of learning and tuning processes for image classification and large language models. For machine learning model in image recognition, various parallelization methods are developed based on different hardware and software scenarios: simple data parallelism, distributed data parallelism, and distributed processing. A detailed description of presented strategies is given, highlighting the challenges and benefits of their application. Furthermore, the impact of different dataset types on the tuning process of large language models is investigated. Experiments show to what extent the task type affects the iteration time in a multi-GPU environment, offering valuable insights into the optimal data utilization strategies to improve model performance. Furthermore, this study leverages the built-in parallelization mechanisms of PyTorch that can facilitate these tasks. Furthermore, performance profiling is incorporated into the study to thoroughly evaluate the impact of memory and communication operations during the training/tuning procedure. Test scenarios are developed and tested with numerous benchmarks on the NVIDIA H100 architecture showing efficiency through selected metrics.
title Efficient allocation of image recognition and LLM tasks on multi-GPU system
topic Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2503.15252