Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Fuqun, Osher, Stanley, Li, Wuchen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preconditioned Regularized Wasserstein Proximal Sampling
by: Tan, Hong Ye, et al.
Published: (2025)
by: Tan, Hong Ye, et al.
Published: (2025)
Tensor train based sampling algorithms for approximating regularized Wasserstein proximal operators
by: Han, Fuqun, et al.
Published: (2024)
by: Han, Fuqun, et al.
Published: (2024)
Accelerated Regularized Wasserstein Proximal Sampling Algorithms
by: Tan, Hong Ye, et al.
Published: (2026)
by: Tan, Hong Ye, et al.
Published: (2026)
Splitting Regularized Wasserstein Proximal Algorithms for Nonsmooth Sampling Problems
by: Han, Fuqun, et al.
Published: (2025)
by: Han, Fuqun, et al.
Published: (2025)
Convergence of Noise-Free Sampling Algorithms with Regularized Wasserstein Proximals
by: Han, Fuqun, et al.
Published: (2024)
by: Han, Fuqun, et al.
Published: (2024)
A Natural Primal-Dual Hybrid Gradient Method for Adversarial Neural Network Training on Solving Partial Differential Equations
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
Score-based Neural Ordinary Differential Equations for Computing Mean Field Control Problems
by: Zhou, Mo, et al.
Published: (2024)
by: Zhou, Mo, et al.
Published: (2024)
Numerical Analysis on Neural Network Projected Schemes for Approximating One Dimensional Wasserstein Gradient Flows
by: Zuo, Xinzhe, et al.
Published: (2024)
by: Zuo, Xinzhe, et al.
Published: (2024)
Gradient-adjusted underdamped Langevin dynamics for sampling
by: Zuo, Xinzhe, et al.
Published: (2024)
by: Zuo, Xinzhe, et al.
Published: (2024)
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
Locally Regularized Sparse Graph by Fast Proximal Gradient Descent
by: Sun, Dongfang, et al.
Published: (2024)
by: Sun, Dongfang, et al.
Published: (2024)
Efficient Computation of Mean field Control based Barycenters from Reaction-Diffusion Systems
by: Vijaywargiya, Arjun, et al.
Published: (2024)
by: Vijaywargiya, Arjun, et al.
Published: (2024)
On Generalization and Regularization via Wasserstein Distributionally Robust Optimization
by: Wu, Qinyu, et al.
Published: (2022)
by: Wu, Qinyu, et al.
Published: (2022)
Numerical analysis of a first-order computational algorithm for reaction-diffusion equations via the primal-dual hybrid gradient method
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
End-to-End Training of High-Dimensional Optimal Control with Implicit Hamiltonians via Jacobian-Free Backpropagation
by: Gelphman, Eric, et al.
Published: (2025)
by: Gelphman, Eric, et al.
Published: (2025)
A Primal-dual hybrid gradient method for solving optimal control problems and the corresponding Hamilton-Jacobi PDEs
by: Meng, Tingwei, et al.
Published: (2024)
by: Meng, Tingwei, et al.
Published: (2024)
Convergence Analysis of the Wasserstein Proximal Algorithm beyond Geodesic Convexity
by: Zhu, Shuailong, et al.
Published: (2025)
by: Zhu, Shuailong, et al.
Published: (2025)
Inexact Proximal Point Algorithms for Zeroth-Order Global Optimization
by: Zhang, Minxin, et al.
Published: (2024)
by: Zhang, Minxin, et al.
Published: (2024)
Zero-Shot Transferable Solution Method for Parametric Optimal Control Problems
by: Li, Xingjian, et al.
Published: (2025)
by: Li, Xingjian, et al.
Published: (2025)
Stability of Transformers under Layer Normalization
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
Proximal Operators of Sorted Nonconvex Penalties
by: Gagneux, Anne, et al.
Published: (2025)
by: Gagneux, Anne, et al.
Published: (2025)
Sparse Deep Learning Models with the $\ell_1$ Regularization
by: Shen, Lixin, et al.
Published: (2024)
by: Shen, Lixin, et al.
Published: (2024)
Operator Splitting for Learning to Predict Equilibria in Convex Games
by: McKenzie, Daniel, et al.
Published: (2021)
by: McKenzie, Daniel, et al.
Published: (2021)
A New Convergence Analysis of Plug-and-Play Proximal Gradient Descent Under Prior Mismatch
by: Xu, Guixian, et al.
Published: (2026)
by: Xu, Guixian, et al.
Published: (2026)
PINS: Proximal Iterations with Sparse Newton and Sinkhorn for Optimal Transport
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
A Proximal Modified Quasi-Newton Method for Nonsmooth Regularized Optimization
by: Diouane, Youssef, et al.
Published: (2024)
by: Diouane, Youssef, et al.
Published: (2024)
A deep learning algorithm for computing mean field control problems via forward-backward score dynamics
by: Zhou, Mo, et al.
Published: (2024)
by: Zhou, Mo, et al.
Published: (2024)
On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization
by: Liang, Ling, et al.
Published: (2024)
by: Liang, Ling, et al.
Published: (2024)
A Unifying View of Anchoring via Operator-Side Tikhonov Regularization
by: Chen, Zihao
Published: (2026)
by: Chen, Zihao
Published: (2026)
SPP-SBL: Space-Power Prior Sparse Bayesian Learning for Block Sparse Recovery
by: Zhang, Yanhao, et al.
Published: (2025)
by: Zhang, Yanhao, et al.
Published: (2025)
Optimal transport natural gradient for statistical manifolds with continuous sample space
by: Chen, Yifan, et al.
Published: (2018)
by: Chen, Yifan, et al.
Published: (2018)
Collaborative Bayesian Optimization via Wasserstein Barycenters
by: Zhan, Donglin, et al.
Published: (2025)
by: Zhan, Donglin, et al.
Published: (2025)
Universal Architectures for the Learning of Polyhedral Norms and Convex Regularizers
by: Unser, Michael, et al.
Published: (2025)
by: Unser, Michael, et al.
Published: (2025)
Flowing Datasets with Wasserstein over Wasserstein Gradient Flows
by: Bonet, Clément, et al.
Published: (2025)
by: Bonet, Clément, et al.
Published: (2025)
Simulating Fokker-Planck equations via mean field control of score-based normalizing flows
by: Zhou, Mo, et al.
Published: (2025)
by: Zhou, Mo, et al.
Published: (2025)
Conditional Sampling via Wasserstein Autoencoders and Triangular Transport
by: Al-Jarrah, Mohammad, et al.
Published: (2026)
by: Al-Jarrah, Mohammad, et al.
Published: (2026)
Smoothing the Edges: Smooth Optimization for Sparse Regularization using Hadamard Overparametrization
by: Kolb, Chris, et al.
Published: (2023)
by: Kolb, Chris, et al.
Published: (2023)
Dynamic Proximal Gradient Algorithms for Schatten-$p$ Quasi-Norm Regularized Problems
by: Shen, Weiping, et al.
Published: (2026)
by: Shen, Weiping, et al.
Published: (2026)
Worst-case generation via minimax optimization in Wasserstein space
by: Cheng, Xiuyuan, et al.
Published: (2025)
by: Cheng, Xiuyuan, et al.
Published: (2025)
Incremental Learning of Sparse Attention Patterns in Transformers
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
Similar Items
-
Preconditioned Regularized Wasserstein Proximal Sampling
by: Tan, Hong Ye, et al.
Published: (2025) -
Tensor train based sampling algorithms for approximating regularized Wasserstein proximal operators
by: Han, Fuqun, et al.
Published: (2024) -
Accelerated Regularized Wasserstein Proximal Sampling Algorithms
by: Tan, Hong Ye, et al.
Published: (2026) -
Splitting Regularized Wasserstein Proximal Algorithms for Nonsmooth Sampling Problems
by: Han, Fuqun, et al.
Published: (2025) -
Convergence of Noise-Free Sampling Algorithms with Regularized Wasserstein Proximals
by: Han, Fuqun, et al.
Published: (2024)