The AI_INFN Platform: Artificial Intelligence Development in the Cloud

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anderlini, Lucio, Bianchini, Giulio, Ciangottini, Diego, Pra, Stefano Dal, Michelotto, Diego, Petrini, Rosa, Spiga, Daniele
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917048803655680
author Anderlini, Lucio
Bianchini, Giulio
Ciangottini, Diego
Pra, Stefano Dal
Michelotto, Diego
Petrini, Rosa
Spiga, Daniele
author_facet Anderlini, Lucio
Bianchini, Giulio
Ciangottini, Diego
Pra, Stefano Dal
Michelotto, Diego
Petrini, Rosa
Spiga, Daniele
contents Machine Learning (ML) is profoundly reshaping the way researchers create, implement, and operate data-intensive software. Its adoption, however, introduces notable challenges for computing infrastructures, particularly when it comes to coordinating access to hardware accelerators across development, testing, and production environments. The INFN initiative AI_INFN (Artificial Intelligence at INFN) seeks to promote the use of ML methods across various INFN research scenarios by offering comprehensive technical support, including access to AI-focused computational resources. Leveraging the INFN Cloud ecosystem and cloud-native technologies, the project emphasizes efficient sharing of accelerator hardware while maintaining the breadth of the Institute's research activities. This contribution describes the deployment and commissioning of a Kubernetes-based platform designed to simplify GPU-powered data analysis workflows and enable their scalable execution on heterogeneous distributed resources. By integrating offloading mechanisms through Virtual Kubelet and the InterLink API, the platform allows workflows to span multiple resource providers, from Worldwide LHC Computing Grid sites to high-performance computing centers like CINECA Leonardo. We will present preliminary benchmarks, functional tests, and case studies, demonstrating both performance and integration outcomes.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22117
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The AI_INFN Platform: Artificial Intelligence Development in the Cloud
Anderlini, Lucio
Bianchini, Giulio
Ciangottini, Diego
Pra, Stefano Dal
Michelotto, Diego
Petrini, Rosa
Spiga, Daniele
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning (ML) is profoundly reshaping the way researchers create, implement, and operate data-intensive software. Its adoption, however, introduces notable challenges for computing infrastructures, particularly when it comes to coordinating access to hardware accelerators across development, testing, and production environments. The INFN initiative AI_INFN (Artificial Intelligence at INFN) seeks to promote the use of ML methods across various INFN research scenarios by offering comprehensive technical support, including access to AI-focused computational resources. Leveraging the INFN Cloud ecosystem and cloud-native technologies, the project emphasizes efficient sharing of accelerator hardware while maintaining the breadth of the Institute's research activities. This contribution describes the deployment and commissioning of a Kubernetes-based platform designed to simplify GPU-powered data analysis workflows and enable their scalable execution on heterogeneous distributed resources. By integrating offloading mechanisms through Virtual Kubelet and the InterLink API, the platform allows workflows to span multiple resource providers, from Worldwide LHC Computing Grid sites to high-performance computing centers like CINECA Leonardo. We will present preliminary benchmarks, functional tests, and case studies, demonstrating both performance and integration outcomes.
title The AI_INFN Platform: Artificial Intelligence Development in the Cloud
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2509.22117