RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Byeongchan, Sehanobish, Arijit, Dubey, Avinava, Oh, Min-hwan, Choromanski, Krzysztof
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910209404829696
author Kim, Byeongchan
Sehanobish, Arijit
Dubey, Avinava
Oh, Min-hwan
Choromanski, Krzysztof
author_facet Kim, Byeongchan
Sehanobish, Arijit
Dubey, Avinava
Oh, Min-hwan
Choromanski, Krzysztof
contents We present a new class of efficient attention mechanisms applying universal 3D Relative Positional Encoding (RPE) methods given by arbitrary integrable modulation functions $f$. They lead to the new class of 3D-Transformer models, called \textit{RelFlexformers}, flexibly integrating those RPEs, and characterized by the $O(L \log L)$ time complexity of the attention computation for the $L$-length input sequences. RelFlexformers builds on the theory of the Non-Uniform Fourier Transform (NU-FFT), naturally generalizing several existing efficient RPE-attention methods from structured settings with tokens homogeneously embedded in unweighted grids into general non-structured heterogeneous scenarios, where tokens' positions are arbitrarily distributed in the corresponding 3D spaces. As such, RelFlexformers can be applied in particular to model point clouds. Our extensive empirical evaluation on a large portfolio of 3D datasets confirms quality improvements provided by the NU-FFT-driven attention modulation techniques in the RelFlexformers.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10706
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings
Kim, Byeongchan
Sehanobish, Arijit
Dubey, Avinava
Oh, Min-hwan
Choromanski, Krzysztof
Machine Learning
We present a new class of efficient attention mechanisms applying universal 3D Relative Positional Encoding (RPE) methods given by arbitrary integrable modulation functions $f$. They lead to the new class of 3D-Transformer models, called \textit{RelFlexformers}, flexibly integrating those RPEs, and characterized by the $O(L \log L)$ time complexity of the attention computation for the $L$-length input sequences. RelFlexformers builds on the theory of the Non-Uniform Fourier Transform (NU-FFT), naturally generalizing several existing efficient RPE-attention methods from structured settings with tokens homogeneously embedded in unweighted grids into general non-structured heterogeneous scenarios, where tokens' positions are arbitrarily distributed in the corresponding 3D spaces. As such, RelFlexformers can be applied in particular to model point clouds. Our extensive empirical evaluation on a large portfolio of 3D datasets confirms quality improvements provided by the NU-FFT-driven attention modulation techniques in the RelFlexformers.
title RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings
topic Machine Learning
url https://arxiv.org/abs/2605.10706