Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lapautre, Nicolas, Marchenko, Maria, Patiño, Carlos Miguel, Zhou, Xin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908490044276736
author Lapautre, Nicolas
Marchenko, Maria
Patiño, Carlos Miguel
Zhou, Xin
author_facet Lapautre, Nicolas
Marchenko, Maria
Patiño, Carlos Miguel
Zhou, Xin
contents Unlocking the potential of transformers on datasets of large physical systems depends on overcoming the quadratic scaling of the attention mechanism. This work explores combining the Erwin architecture with the Native Sparse Attention (NSA) mechanism to improve the efficiency and receptive field of transformer models for large-scale physical systems, addressing the challenge of quadratic attention complexity. We adapt the NSA mechanism for non-sequential data, implement the Erwin NSA model, and evaluate it on three datasets from the physical sciences -- cosmology simulations, molecular dynamics, and air pressure modeling -- achieving performance that matches or exceeds that of the original Erwin model. Additionally, we reproduce the experimental results from the Erwin paper to validate their implementation.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10758
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
Lapautre, Nicolas
Marchenko, Maria
Patiño, Carlos Miguel
Zhou, Xin
Machine Learning
Artificial Intelligence
Unlocking the potential of transformers on datasets of large physical systems depends on overcoming the quadratic scaling of the attention mechanism. This work explores combining the Erwin architecture with the Native Sparse Attention (NSA) mechanism to improve the efficiency and receptive field of transformer models for large-scale physical systems, addressing the challenge of quadratic attention complexity. We adapt the NSA mechanism for non-sequential data, implement the Erwin NSA model, and evaluate it on three datasets from the physical sciences -- cosmology simulations, molecular dynamics, and air pressure modeling -- achieving performance that matches or exceeds that of the original Erwin model. Additionally, we reproduce the experimental results from the Erwin paper to validate their implementation.
title Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.10758