SkimROOT: Accelerating LHC Data Filtering with Near-Storage Processing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Batsoyol, Narangerelt, Guiang, Jonathan, Davila, Diego, Arora, Aashay, Chang, Philip, Würthwein, Frank, Swanson, Steven
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909638862045184
author Batsoyol, Narangerelt
Guiang, Jonathan
Davila, Diego
Arora, Aashay
Chang, Philip
Würthwein, Frank
Swanson, Steven
author_facet Batsoyol, Narangerelt
Guiang, Jonathan
Davila, Diego
Arora, Aashay
Chang, Philip
Würthwein, Frank
Swanson, Steven
contents Data analysis in high-energy physics (HEP) begins with data reduction, where vast datasets are filtered to extract relevant events. At the Large Hadron Collider (LHC), this process is bottlenecked by slow data transfers between storage and compute nodes. To address this, we introduce SkimROOT, a near-data filtering system leveraging Data Processing Units (DPUs) to accelerate LHC data analysis. By performing filtering directly on storage servers and returning only the relevant data, SkimROOT minimizes data movement and reduces processing delays. Our prototype demonstrates significant efficiency gains, achieving a 44.3$\times$ performance improvement, paving the way for faster physics discoveries.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SkimROOT: Accelerating LHC Data Filtering with Near-Storage Processing
Batsoyol, Narangerelt
Guiang, Jonathan
Davila, Diego
Arora, Aashay
Chang, Philip
Würthwein, Frank
Swanson, Steven
Distributed, Parallel, and Cluster Computing
Data analysis in high-energy physics (HEP) begins with data reduction, where vast datasets are filtered to extract relevant events. At the Large Hadron Collider (LHC), this process is bottlenecked by slow data transfers between storage and compute nodes. To address this, we introduce SkimROOT, a near-data filtering system leveraging Data Processing Units (DPUs) to accelerate LHC data analysis. By performing filtering directly on storage servers and returning only the relevant data, SkimROOT minimizes data movement and reduces processing delays. Our prototype demonstrates significant efficiency gains, achieving a 44.3$\times$ performance improvement, paving the way for faster physics discoveries.
title SkimROOT: Accelerating LHC Data Filtering with Near-Storage Processing
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2506.04507