Sinkhorn doubly stochastic attention rank decay analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Lapenna, Michela, Fioresi, Rita, Gharesifard, Bahman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Topology of Neural Network Superlevel Sets
by: Gharesifard, Bahman
Published: (2026)
by: Gharesifard, Bahman
Published: (2026)
Universal Approximation Power of Deep Residual Neural Networks via Nonlinear Control Theory
by: Tabuada, Paulo, et al.
Published: (2020)
by: Tabuada, Paulo, et al.
Published: (2020)
Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens
by: Moghaddam, Amirreza Neshaei, et al.
Published: (2024)
by: Moghaddam, Amirreza Neshaei, et al.
Published: (2024)
Sample Complexity of Linear Quadratic Regulator Without Initial Stability
by: Moghaddam, Amirreza Neshaei, et al.
Published: (2025)
by: Moghaddam, Amirreza Neshaei, et al.
Published: (2025)
Feature-aligned N-BEATS with Sinkhorn divergence
by: Lee, Joonhun, et al.
Published: (2023)
by: Lee, Joonhun, et al.
Published: (2023)
Nonlinear Non-Gaussian Density Steering with Input and Noise Channel Mismatch: Sinkhorn with Memory for Solving the Control-affine Schrödinger Bridge Problem
by: Bondar, Georgiy A., et al.
Published: (2026)
by: Bondar, Georgiy A., et al.
Published: (2026)
Sinkhorn Based Associative Memory Retrieval Using Spherical Hellinger Kantorovich Dynamics
by: Mustafi, Aratrika, et al.
Published: (2026)
by: Mustafi, Aratrika, et al.
Published: (2026)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Localmax dynamics for attention in transformers and its asymptotic behavior
by: Cimetière, Henri, et al.
Published: (2025)
by: Cimetière, Henri, et al.
Published: (2025)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
by: Xu, Conglong, et al.
Published: (2025)
by: Xu, Conglong, et al.
Published: (2025)
A space-decoupling framework for optimization on bounded-rank matrices with orthogonally invariant constraints
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
Fast solution to the fair ranking problem using the Sinkhorn algorithm
by: Uehara, Yuki, et al.
Published: (2024)
by: Uehara, Yuki, et al.
Published: (2024)
Sinkhorn Distributionally Robust Optimization
by: Wang, Jie, et al.
Published: (2021)
by: Wang, Jie, et al.
Published: (2021)
On Reward-Balancing Methods for Reinforcement Learning
by: Baroncini, Simone, et al.
Published: (2026)
by: Baroncini, Simone, et al.
Published: (2026)
Inferring Global Exponential Stability Properties using Lie-bracket Approximations
by: Weber, Marc, et al.
Published: (2024)
by: Weber, Marc, et al.
Published: (2024)
Flexible-step Model Predictive Control based on Generalized Lyapunov Functions
by: Fürnsinn, Annika, et al.
Published: (2022)
by: Fürnsinn, Annika, et al.
Published: (2022)
Flexible-step MPC for Switched Linear Systems with No Quadratic Common Lyapunov Function
by: Fürnsinn, Annika, et al.
Published: (2024)
by: Fürnsinn, Annika, et al.
Published: (2024)
Accelerating Sinkhorn Algorithm with Sparse Newton Iterations
by: Tang, Xun, et al.
Published: (2024)
by: Tang, Xun, et al.
Published: (2024)
On Sinkhorn's Algorithm and Choice Modeling
by: Qu, Zhaonan, et al.
Published: (2023)
by: Qu, Zhaonan, et al.
Published: (2023)
A Sinkhorn-type Algorithm for Constrained Optimal Transport
by: Tang, Xun, et al.
Published: (2024)
by: Tang, Xun, et al.
Published: (2024)
From Schrodinger Bridge to Optimal Transport over Sub-Riemannian Manifolds
by: Adu, Daniel Owusu, et al.
Published: (2026)
by: Adu, Daniel Owusu, et al.
Published: (2026)
Annealed Sinkhorn for Optimal Transport: convergence, regularization path and debiasing
by: Chizat, Lénaïc
Published: (2024)
by: Chizat, Lénaïc
Published: (2024)
PINS: Proximal Iterations with Sparse Newton and Sinkhorn for Optimal Transport
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
An Approximate Ascent Approach To Prove Convergence of PPO
by: Doering, Leif, et al.
Published: (2026)
by: Doering, Leif, et al.
Published: (2026)
On a Gradient Approach to Chebyshev Center Problems with Applications to Function Learning
by: Raghuvanshi, Abhinav, et al.
Published: (2026)
by: Raghuvanshi, Abhinav, et al.
Published: (2026)
Delightful Distributed Policy Gradient
by: Osband, Ian
Published: (2026)
by: Osband, Ian
Published: (2026)
Optimal Pattern Detection Tree for Symbolic Rule-Based Classification
by: Hong, Young-Chae, et al.
Published: (2026)
by: Hong, Young-Chae, et al.
Published: (2026)
Feature Starvation as Geometric Instability in Sparse Autoencoders
by: Chaudhry, Faris, et al.
Published: (2026)
by: Chaudhry, Faris, et al.
Published: (2026)
Predictive and Prescriptive AI toward Optimizing Wildfire Suppression
by: Boussioux, Leonard, et al.
Published: (2026)
by: Boussioux, Leonard, et al.
Published: (2026)
$γ$-weakly $θ$-up-concavity: A Unified Framework for Non-Convex Optimization Beyond DR-Submodular and OSS Functions
by: Pedramfar, Mohammad, et al.
Published: (2026)
by: Pedramfar, Mohammad, et al.
Published: (2026)
From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
by: Lin, Jianghao, et al.
Published: (2026)
by: Lin, Jianghao, et al.
Published: (2026)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
by: Molodtsov, Gleb, et al.
Published: (2026)
by: Molodtsov, Gleb, et al.
Published: (2026)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Budget-aware Auto Optimizer Configurator
by: Liu, Kang, et al.
Published: (2026)
by: Liu, Kang, et al.
Published: (2026)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling
by: Hao, Yongchang, et al.
Published: (2026)
by: Hao, Yongchang, et al.
Published: (2026)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Delightful Policy Gradient
by: Osband, Ian
Published: (2026)
by: Osband, Ian
Published: (2026)
Similar Items
-
On the Topology of Neural Network Superlevel Sets
by: Gharesifard, Bahman
Published: (2026) -
Universal Approximation Power of Deep Residual Neural Networks via Nonlinear Control Theory
by: Tabuada, Paulo, et al.
Published: (2020) -
Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens
by: Moghaddam, Amirreza Neshaei, et al.
Published: (2024) -
Sample Complexity of Linear Quadratic Regulator Without Initial Stability
by: Moghaddam, Amirreza Neshaei, et al.
Published: (2025) -
Feature-aligned N-BEATS with Sinkhorn divergence
by: Lee, Joonhun, et al.
Published: (2023)