A universal vision transformer for fast calorimeter simulations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Favaro, Luigi, Giammanco, Andrea, Krause, Claudius
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911714975416320
author Favaro, Luigi
Giammanco, Andrea
Krause, Claudius
author_facet Favaro, Luigi
Giammanco, Andrea
Krause, Claudius
contents The high-dimensional complex nature of detectors makes fast calorimeter simulations a prime application for modern generative machine learning. Vision transformers (ViTs) can emulate the Geant4 response with unmatched accuracy and are not limited to regular geometries. Starting from the CaloDREAM architecture, we demonstrate the robustness and scalability of ViTs on regular and irregular geometries, and multiple detectors. Our results show that ViTs generate electromagnetic and hadronic showers with minimal deviations from Geant4 in multiple evaluation metrics, while maintaining the generation time in the $\mathcal{O}(10-100)$ ms on a single GPU. Furthermore, we show that pretraining on a large dataset and fine-tuning on the target geometry leads to reduced training costs and higher data efficiency, or altogether improves the fidelity of generated showers.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05289
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A universal vision transformer for fast calorimeter simulations
Favaro, Luigi
Giammanco, Andrea
Krause, Claudius
High Energy Physics - Phenomenology
Machine Learning
High Energy Physics - Experiment
Instrumentation and Detectors
The high-dimensional complex nature of detectors makes fast calorimeter simulations a prime application for modern generative machine learning. Vision transformers (ViTs) can emulate the Geant4 response with unmatched accuracy and are not limited to regular geometries. Starting from the CaloDREAM architecture, we demonstrate the robustness and scalability of ViTs on regular and irregular geometries, and multiple detectors. Our results show that ViTs generate electromagnetic and hadronic showers with minimal deviations from Geant4 in multiple evaluation metrics, while maintaining the generation time in the $\mathcal{O}(10-100)$ ms on a single GPU. Furthermore, we show that pretraining on a large dataset and fine-tuning on the target geometry leads to reduced training costs and higher data efficiency, or altogether improves the fidelity of generated showers.
title A universal vision transformer for fast calorimeter simulations
topic High Energy Physics - Phenomenology
Machine Learning
High Energy Physics - Experiment
Instrumentation and Detectors
url https://arxiv.org/abs/2601.05289