How Smoothing is N-simplicial Attention?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dussolle, Alexandre, Liò, Pietro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918252640206848
author Dussolle, Alexandre
Liò, Pietro
author_facet Dussolle, Alexandre
Liò, Pietro
contents Going from pure Multilayer Perceptron (MLP) to a learnable graph message-passing mechanism at each layer has been foundational to state-of-the-art results, despite the computational trade-off (e.g. GATs or Transformers). To go a step further, in this work, we introduce N-simplicial attention, going from pairwise token similarity to higher-order interactions, and adapt it for Rotary Position Embeddings (RoPE). To help manage the increased complexity, we propose a cost-effective simplex selection enabling the model to focus its computation load onto the more task-sensitive interactions. Beyond these core mechanisms, we study how smoothing N-simplicial attention is by deriving a Lipschitz upper-bound and by demonstrating that by itself it also suffers from over-smoothing, despite opening the attention message-passing to higher-order interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15600
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Smoothing is N-simplicial Attention?
Dussolle, Alexandre
Liò, Pietro
Machine Learning
Artificial Intelligence
Going from pure Multilayer Perceptron (MLP) to a learnable graph message-passing mechanism at each layer has been foundational to state-of-the-art results, despite the computational trade-off (e.g. GATs or Transformers). To go a step further, in this work, we introduce N-simplicial attention, going from pairwise token similarity to higher-order interactions, and adapt it for Rotary Position Embeddings (RoPE). To help manage the increased complexity, we propose a cost-effective simplex selection enabling the model to focus its computation load onto the more task-sensitive interactions. Beyond these core mechanisms, we study how smoothing N-simplicial attention is by deriving a Lipschitz upper-bound and by demonstrating that by itself it also suffers from over-smoothing, despite opening the attention message-passing to higher-order interactions.
title How Smoothing is N-simplicial Attention?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.15600