Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Fuqun, Osher, Stanley, Li, Wuchen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917024363446272
author Han, Fuqun
Osher, Stanley
Li, Wuchen
author_facet Han, Fuqun
Osher, Stanley
Li, Wuchen
contents In this work, we propose a sparse transformer architecture that incorporates prior information about the underlying data distribution directly into the transformer structure of the neural network. The design of the model is motivated by a special optimal transport problem, namely the regularized Wasserstein proximal operator, which admits a closed-form solution and turns out to be a special representation of transformer architectures. Compared with classical flow-based models, the proposed approach improves the convexity properties of the optimization problem and promotes sparsity in the generated samples. Through both theoretical analysis and numerical experiments, including applications in generative modeling and Bayesian inverse problems, we demonstrate that the sparse transformer achieves higher accuracy and faster convergence to the target distribution than classical neural ODE-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16356
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior
Han, Fuqun
Osher, Stanley
Li, Wuchen
Machine Learning
Optimization and Control
In this work, we propose a sparse transformer architecture that incorporates prior information about the underlying data distribution directly into the transformer structure of the neural network. The design of the model is motivated by a special optimal transport problem, namely the regularized Wasserstein proximal operator, which admits a closed-form solution and turns out to be a special representation of transformer architectures. Compared with classical flow-based models, the proposed approach improves the convexity properties of the optimization problem and promotes sparsity in the generated samples. Through both theoretical analysis and numerical experiments, including applications in generative modeling and Bayesian inverse problems, we demonstrate that the sparse transformer achieves higher accuracy and faster convergence to the target distribution than classical neural ODE-based methods.
title Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2510.16356