Transformers for Tabular Data: A Training Perspective of Self-Attention via Optimal Transport
Fuente:
arXiv
Saved in:
| Main Authors: | Quadrio, Alessandro, Candelieri, Antonio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weighted Wasserstein Barycenter of Gaussian Processes for exotic Bayesian Optimization tasks
by: Candelieri, Antonio, et al.
Published: (2026)
by: Candelieri, Antonio, et al.
Published: (2026)
Calibrating Tabular Anomaly Detection via Optimal Transport
by: Ye, Hangting, et al.
Published: (2026)
by: Ye, Hangting, et al.
Published: (2026)
Information Theoretic Bayesian Optimization over the Probability Simplex
by: Pavesi, Federico, et al.
Published: (2026)
by: Pavesi, Federico, et al.
Published: (2026)
Multi-Layer Attention-Based Explainability via Transformers for Tabular Data
by: Gavito, Andrea Treviño, et al.
Published: (2023)
by: Gavito, Andrea Treviño, et al.
Published: (2023)
Rethinking Pre-Training in Tabular Data: A Neighborhood Embedding Perspective
by: Ye, Han-Jia, et al.
Published: (2023)
by: Ye, Han-Jia, et al.
Published: (2023)
CAST: Cluster-Aware Self-Training for Tabular Data via Reliable Confidence
by: Kim, Minwook, et al.
Published: (2023)
by: Kim, Minwook, et al.
Published: (2023)
Wasserstein Barycenter Gaussian Process based Bayesian Optimization
by: Candelieri, Antonio, et al.
Published: (2025)
by: Candelieri, Antonio, et al.
Published: (2025)
Fine-grained Attention in Hierarchical Transformers for Tabular Time-series
by: Azorin, Raphael, et al.
Published: (2024)
by: Azorin, Raphael, et al.
Published: (2024)
Gromov-Wasserstein and optimal transport: from assignment problems to probabilistic numeric
by: Seyedi, Iman, et al.
Published: (2025)
by: Seyedi, Iman, et al.
Published: (2025)
A Systematic Evaluation of Generative Models on Tabular Transportation Data
by: Wang, Chengen, et al.
Published: (2025)
by: Wang, Chengen, et al.
Published: (2025)
Attention versus Contrastive Learning of Tabular Data -- A Data-centric Benchmarking
by: Rabbani, Shourav B., et al.
Published: (2024)
by: Rabbani, Shourav B., et al.
Published: (2024)
Transformers with Stochastic Competition for Tabular Data Modelling
by: Voskou, Andreas, et al.
Published: (2024)
by: Voskou, Andreas, et al.
Published: (2024)
Transformer Fusion with Optimal Transport
by: Imfeld, Moritz, et al.
Published: (2023)
by: Imfeld, Moritz, et al.
Published: (2023)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Interpretable Feature Interaction via Statistical Self-supervised Learning on Tabular Data
by: Zhang, Xiaochen, et al.
Published: (2025)
by: Zhang, Xiaochen, et al.
Published: (2025)
Towards a Relationship-Aware Transformer for Tabular Data
by: Konstantinov, Andrei V., et al.
Published: (2025)
by: Konstantinov, Andrei V., et al.
Published: (2025)
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
by: Shahbazi, Ashkan, et al.
Published: (2025)
by: Shahbazi, Ashkan, et al.
Published: (2025)
Optimal Transport-based Permutation-Invariant Bayesian Optimization of Offshore Wind Farm Layouts
by: Candelieri, Antonio, et al.
Published: (2026)
by: Candelieri, Antonio, et al.
Published: (2026)
TabNSA: Native Sparse Attention for Efficient Tabular Data Learning
by: Eslamian, Ali, et al.
Published: (2025)
by: Eslamian, Ali, et al.
Published: (2025)
Domain Adaptation and Entanglement: an Optimal Transport Perspective
by: Koç, Okan, et al.
Published: (2025)
by: Koç, Okan, et al.
Published: (2025)
Utilizing Training Data to Improve LLM Reasoning for Tabular Understanding
by: Gao, Chufan, et al.
Published: (2025)
by: Gao, Chufan, et al.
Published: (2025)
One Transformer for All Time Series: Representing and Training with Time-Dependent Heterogeneous Tabular Data
by: Luetto, Simone, et al.
Published: (2023)
by: Luetto, Simone, et al.
Published: (2023)
Self-Improving Tabular Language Models via Iterative Reward-Guided Post-Training
by: Long, Yunbo, et al.
Published: (2026)
by: Long, Yunbo, et al.
Published: (2026)
Reliable Pseudo-labeling via Optimal Transport with Attention for Short Text Clustering
by: Yao, Zhihao
Published: (2025)
by: Yao, Zhihao
Published: (2025)
A Data-Centric Perspective on Evaluating Machine Learning Models for Tabular Data
by: Tschalzev, Andrej, et al.
Published: (2024)
by: Tschalzev, Andrej, et al.
Published: (2024)
A Closer Look on Memorization in Tabular Diffusion Model: A Data-Centric Perspective
by: Fang, Zhengyu, et al.
Published: (2025)
by: Fang, Zhengyu, et al.
Published: (2025)
Self-Supervision Improves Diffusion Models for Tabular Data Imputation
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular Data
by: Thimonier, Hugo, et al.
Published: (2024)
by: Thimonier, Hugo, et al.
Published: (2024)
Multi-branch of Attention Yields Accurate Results for Tabular Data
by: Li, Xuechen, et al.
Published: (2025)
by: Li, Xuechen, et al.
Published: (2025)
Deep Learning with Tabular Data: A Self-supervised Approach
by: Vyas, Tirth Kiranbhai
Published: (2024)
by: Vyas, Tirth Kiranbhai
Published: (2024)
Diffusion Transformers for Tabular Data Time Series Generation
by: Garuti, Fabrizio, et al.
Published: (2025)
by: Garuti, Fabrizio, et al.
Published: (2025)
Backdoor Attacks on Transformers for Tabular Data: An Empirical Study
by: Pleiter, Bart, et al.
Published: (2023)
by: Pleiter, Bart, et al.
Published: (2023)
GeoAggregator: An Efficient Transformer Model for Geo-Spatial Tabular Data
by: Deng, Rui, et al.
Published: (2025)
by: Deng, Rui, et al.
Published: (2025)
Fine-tuned In-Context Learning Transformers are Excellent Tabular Data Classifiers
by: Breejen, Felix den, et al.
Published: (2024)
by: Breejen, Felix den, et al.
Published: (2024)
A Sobering Look at Tabular Data Generation via Probabilistic Circuits
by: Scassola, Davide, et al.
Published: (2026)
by: Scassola, Davide, et al.
Published: (2026)
Entropic Riemannian Neural Optimal Transport
by: Micheli, Alessandro, et al.
Published: (2026)
by: Micheli, Alessandro, et al.
Published: (2026)
Attention-Based Deep Learning for Early Parkinson's Disease Detection with Tabular Biomedical Data
by: Oseni, Olamide Samuel, et al.
Published: (2026)
by: Oseni, Olamide Samuel, et al.
Published: (2026)
Riemannian Neural Optimal Transport
by: Micheli, Alessandro, et al.
Published: (2026)
by: Micheli, Alessandro, et al.
Published: (2026)
Weight-Informed Self-Explaining Clustering for Mixed-Type Tabular Data
by: Li, Lehao, et al.
Published: (2026)
by: Li, Lehao, et al.
Published: (2026)
Scaled-Dot-Product Attention as One-Sided Entropic Optimal Transport
by: Litman, Elon
Published: (2025)
by: Litman, Elon
Published: (2025)
Similar Items
-
Weighted Wasserstein Barycenter of Gaussian Processes for exotic Bayesian Optimization tasks
by: Candelieri, Antonio, et al.
Published: (2026) -
Calibrating Tabular Anomaly Detection via Optimal Transport
by: Ye, Hangting, et al.
Published: (2026) -
Information Theoretic Bayesian Optimization over the Probability Simplex
by: Pavesi, Federico, et al.
Published: (2026) -
Multi-Layer Attention-Based Explainability via Transformers for Tabular Data
by: Gavito, Andrea Treviño, et al.
Published: (2023) -
Rethinking Pre-Training in Tabular Data: A Neighborhood Embedding Perspective
by: Ye, Han-Jia, et al.
Published: (2023)