LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
Fuente:
arXiv
Guardado en:
| Autores principales: | Shahbazi, Ashkan, Thrash, Chayne, Bai, Yikun, Hamm, Keaton, NaderiAlizadeh, Navid, Kolouri, Soheil |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LUNA: Linear Universal Neural Attention with Generalization Guarantees
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
ESPFormer: Doubly-Stochastic Attention with Expected Sliced Transport Plans
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
por: Thrash, Chayne, et al.
Publicado: (2026)
por: Thrash, Chayne, et al.
Publicado: (2026)
Constrained Sliced Wasserstein Embedding
por: NaderiAlizadeh, Navid, et al.
Publicado: (2025)
por: NaderiAlizadeh, Navid, et al.
Publicado: (2025)
Low-Rank Prehab: Preparing Neural Networks for SVD Compression
por: Qin, Haoran, et al.
Publicado: (2025)
por: Qin, Haoran, et al.
Publicado: (2025)
Efficient Transferable Optimal Transport via Min-Sliced Transport Plans
por: Liu, Xinran, et al.
Publicado: (2025)
por: Liu, Xinran, et al.
Publicado: (2025)
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
por: Abbasi, Ali, et al.
Publicado: (2026)
por: Abbasi, Ali, et al.
Publicado: (2026)
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
por: Abbasi, Ali, et al.
Publicado: (2026)
por: Abbasi, Ali, et al.
Publicado: (2026)
Stochastic Unrolled Federated Learning
por: Hadou, Samar, et al.
Publicado: (2023)
por: Hadou, Samar, et al.
Publicado: (2023)
Robust Stochastically-Descending Unrolled Networks
por: Hadou, Samar, et al.
Publicado: (2023)
por: Hadou, Samar, et al.
Publicado: (2023)
Linear Spherical Sliced Optimal Transport: A Fast Metric for Comparing Spherical Data
por: Liu, Xinran, et al.
Publicado: (2024)
por: Liu, Xinran, et al.
Publicado: (2024)
Expected Sliced Transport Plans
por: Liu, Xinran, et al.
Publicado: (2024)
por: Liu, Xinran, et al.
Publicado: (2024)
Understanding Learning with Sliced-Wasserstein Requires Rethinking Informative Slices
por: Tran, Huy, et al.
Publicado: (2024)
por: Tran, Huy, et al.
Publicado: (2024)
Linear Optimal Partial Transport Embedding
por: Bai, Yikun, et al.
Publicado: (2023)
por: Bai, Yikun, et al.
Publicado: (2023)
Opportunistic Routing in Wireless Communications via Learnable State-Augmented Policies
por: Das, Sourajit, et al.
Publicado: (2025)
por: Das, Sourajit, et al.
Publicado: (2025)
Learning State-Augmented Policies for Information Routing in Communication Networks
por: Das, Sourajit, et al.
Publicado: (2023)
por: Das, Sourajit, et al.
Publicado: (2023)
Min Generalized Sliced Gromov Wasserstein: A Scalable Path to Gromov Wasserstein
por: Shahbazi, Ashkan, et al.
Publicado: (2026)
por: Shahbazi, Ashkan, et al.
Publicado: (2026)
One Category One Prompt: Dataset Distillation using Diffusion Models
por: Abbasi, Ali, et al.
Publicado: (2024)
por: Abbasi, Ali, et al.
Publicado: (2024)
Neighbor Embeddings Using Unbalanced Optimal Transport Metrics
por: Rana, Muhammad, et al.
Publicado: (2025)
por: Rana, Muhammad, et al.
Publicado: (2025)
MCNC: Manifold-Constrained Reparameterization for Neural Compression
por: Thrash, Chayne, et al.
Publicado: (2024)
por: Thrash, Chayne, et al.
Publicado: (2024)
Primal Dual Continual Learning: Balancing Stability and Plasticity through Adaptive Memory Allocation
por: Elenter, Juan, et al.
Publicado: (2023)
por: Elenter, Juan, et al.
Publicado: (2023)
Stereographic Spherical Sliced Wasserstein Distances
por: Tran, Huy, et al.
Publicado: (2024)
por: Tran, Huy, et al.
Publicado: (2024)
Decentralized Learning Strategies for Estimation Error Minimization with Graph Neural Networks
por: Chen, Xingran, et al.
Publicado: (2026)
por: Chen, Xingran, et al.
Publicado: (2026)
Fast State-Augmented Learning for Wireless Resource Allocation with Dual Variable Regression
por: Uslu, Yigit Berkay, et al.
Publicado: (2025)
por: Uslu, Yigit Berkay, et al.
Publicado: (2025)
Transferable Graphical MARL for Real-Time Estimation in Dynamic Wireless Networks
por: Chen, Xingran, et al.
Publicado: (2024)
por: Chen, Xingran, et al.
Publicado: (2024)
Learning to Slice Wi-Fi Networks: A State-Augmented Primal-Dual Approach
por: Uslu, Yiğit Berkay, et al.
Publicado: (2024)
por: Uslu, Yiğit Berkay, et al.
Publicado: (2024)
Equivariant vs. Invariant Layers: A Comparison of Backbone and Pooling for Point Cloud Classification
por: Kothapalli, Abihith, et al.
Publicado: (2023)
por: Kothapalli, Abihith, et al.
Publicado: (2023)
Linear Partial Gromov-Wasserstein Embedding
por: Bai, Yikun, et al.
Publicado: (2024)
por: Bai, Yikun, et al.
Publicado: (2024)
Vector-Quantized Soft Label Compression for Dataset Distillation
por: Abbasi, Ali, et al.
Publicado: (2026)
por: Abbasi, Ali, et al.
Publicado: (2026)
OT-MeanFlow3D: Bridging Optimal Transport and Meanflow for Efficient 3D Point Cloud Generation
por: Akbari, Elaheh, et al.
Publicado: (2025)
por: Akbari, Elaheh, et al.
Publicado: (2025)
Sinkhorn-Drifting Generative Models
por: He, Ping, et al.
Publicado: (2026)
por: He, Ping, et al.
Publicado: (2026)
Wireless Link Scheduling with State-Augmented Graph Neural Networks
por: Camargo, Romina Garcia, et al.
Publicado: (2025)
por: Camargo, Romina Garcia, et al.
Publicado: (2025)
On Wasserstein distances for affine transformations of random vectors
por: Hamm, Keaton, et al.
Publicado: (2023)
por: Hamm, Keaton, et al.
Publicado: (2023)
Fused Partial Gromov-Wasserstein for Structured Objects
por: Bai, Yikun, et al.
Publicado: (2025)
por: Bai, Yikun, et al.
Publicado: (2025)
Wasserstein approximation schemes based on Voronoi partitions
por: Hamm, Keaton, et al.
Publicado: (2023)
por: Hamm, Keaton, et al.
Publicado: (2023)
Partial Gromov-Wasserstein Metric
por: Bai, Yikun, et al.
Publicado: (2024)
por: Bai, Yikun, et al.
Publicado: (2024)
Physics informed cell representations for variational formulation of multiscale problems
por: Gao, Yuxiang, et al.
Publicado: (2024)
por: Gao, Yuxiang, et al.
Publicado: (2024)
Reinforcement Learning-Based Optimization of CT Acquisition and Reconstruction Parameters Through Virtual Imaging Trials
por: Fenwick, David, et al.
Publicado: (2025)
por: Fenwick, David, et al.
Publicado: (2025)
Transport Clustering: Solving Low-Rank Optimal Transport via Clustering
por: Schmidt, Henri, et al.
Publicado: (2026)
por: Schmidt, Henri, et al.
Publicado: (2026)
EMPEROR: Efficient Moment-Preserving Representation of Distributions
por: Liu, Xinran, et al.
Publicado: (2025)
por: Liu, Xinran, et al.
Publicado: (2025)
Ejemplares similares
-
LUNA: Linear Universal Neural Attention with Generalization Guarantees
por: Shahbazi, Ashkan, et al.
Publicado: (2025) -
ESPFormer: Doubly-Stochastic Attention with Expected Sliced Transport Plans
por: Shahbazi, Ashkan, et al.
Publicado: (2025) -
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
por: Thrash, Chayne, et al.
Publicado: (2026) -
Constrained Sliced Wasserstein Embedding
por: NaderiAlizadeh, Navid, et al.
Publicado: (2025) -
Low-Rank Prehab: Preparing Neural Networks for SVD Compression
por: Qin, Haoran, et al.
Publicado: (2025)