Transformer Fusion with Optimal Transport
Fuente:
arXiv
Saved in:
| Main Authors: | Imfeld, Moritz, Graldi, Jacopo, Giordano, Marco, Hofmann, Thomas, Anagnostidis, Sotiris, Singh, Sidak Pal |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Meta-Pruning via Optimal Transport
by: Theus, Alexander, et al.
Published: (2024)
by: Theus, Alexander, et al.
Published: (2024)
Model Fusion via Retrofitting
by: Luenam, Phoomraphee, et al.
Published: (2025)
by: Luenam, Phoomraphee, et al.
Published: (2025)
Generalized Linear Mode Connectivity for Transformers
by: Theus, Alexander, et al.
Published: (2025)
by: Theus, Alexander, et al.
Published: (2025)
Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
by: Anagnostidis, Sotiris, et al.
Published: (2023)
by: Anagnostidis, Sotiris, et al.
Published: (2023)
How Susceptible are LLMs to Influence in Prompts?
by: Anagnostidis, Sotiris, et al.
Published: (2024)
by: Anagnostidis, Sotiris, et al.
Published: (2024)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
by: Ormaniec, Weronika, et al.
Published: (2024)
by: Ormaniec, Weronika, et al.
Published: (2024)
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
by: Khromov, Grigory, et al.
Published: (2023)
by: Khromov, Grigory, et al.
Published: (2023)
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025)
by: Zakarin, Daniyar, et al.
Published: (2025)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
by: Anagnostidis, Sotiris, et al.
Published: (2023)
by: Anagnostidis, Sotiris, et al.
Published: (2023)
The Importance of Being Lazy: Scaling Limits of Continual Learning
by: Graldi, Jacopo, et al.
Published: (2025)
by: Graldi, Jacopo, et al.
Published: (2025)
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
by: Singh, Sidak Pal, et al.
Published: (2024)
by: Singh, Sidak Pal, et al.
Published: (2024)
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Local vs Global continual learning
by: Lanzillotta, Giulia, et al.
Published: (2024)
by: Lanzillotta, Giulia, et al.
Published: (2024)
Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers
by: Bill, Eric Tillman, et al.
Published: (2025)
by: Bill, Eric Tillman, et al.
Published: (2025)
Landscaping Linear Mode Connectivity
by: Singh, Sidak Pal, et al.
Published: (2024)
by: Singh, Sidak Pal, et al.
Published: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025)
by: Singh, Sidak Pal, et al.
Published: (2025)
FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute
by: Anagnostidis, Sotiris, et al.
Published: (2025)
by: Anagnostidis, Sotiris, et al.
Published: (2025)
Fusion of Graph Neural Networks via Optimal Transport
by: Ormaniec, Weronika, et al.
Published: (2025)
by: Ormaniec, Weronika, et al.
Published: (2025)
Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment
by: Bachmann, Gregor, et al.
Published: (2025)
by: Bachmann, Gregor, et al.
Published: (2025)
Zero-Shot Offline Imitation Learning via Optimal Transport
by: Rupf, Thomas, et al.
Published: (2024)
by: Rupf, Thomas, et al.
Published: (2024)
Wasserstein Wormhole: Scalable Optimal Transport Distance with Transformers
by: Haviv, Doron, et al.
Published: (2024)
by: Haviv, Doron, et al.
Published: (2024)
Slicing Wasserstein Over Wasserstein Via Functional Optimal Transport
by: Piening, Moritz, et al.
Published: (2025)
by: Piening, Moritz, et al.
Published: (2025)
Multivariate Conformal Prediction using Optimal Transport
by: Klein, Michal, et al.
Published: (2025)
by: Klein, Michal, et al.
Published: (2025)
Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization
by: Thomas, Rahul Krishna, et al.
Published: (2025)
by: Thomas, Rahul Krishna, et al.
Published: (2025)
Progressive Entropic Optimal Transport Solvers
by: Kassraie, Parnian, et al.
Published: (2024)
by: Kassraie, Parnian, et al.
Published: (2024)
MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation
by: Yang, Han, et al.
Published: (2024)
by: Yang, Han, et al.
Published: (2024)
Meta-Learning for Unsupervised Outlier Detection with Optimal Transport
by: Singh, Prabhant, et al.
Published: (2022)
by: Singh, Prabhant, et al.
Published: (2022)
OT-Transformer: A Continuous-time Transformer Architecture with Optimal Transport Regularization
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
Hybrid Generative Modeling for Incomplete Physics: Deep Grey-Box Meets Optimal Transport
by: Singh, Gurjeet Sangra, et al.
Published: (2025)
by: Singh, Gurjeet Sangra, et al.
Published: (2025)
Transformers for Tabular Data: A Training Perspective of Self-Attention via Optimal Transport
by: Quadrio, Alessandro, et al.
Published: (2025)
by: Quadrio, Alessandro, et al.
Published: (2025)
Conditional Variable Flow Matching: Transforming Conditional Densities with Amortized Conditional Optimal Transport
by: Generale, Adam P., et al.
Published: (2024)
by: Generale, Adam P., et al.
Published: (2024)
Anchor Space Optimal Transport as a Fast Solution to Multiple Optimal Transport Problems
by: Huang, Jianming, et al.
Published: (2023)
by: Huang, Jianming, et al.
Published: (2023)
A Language Model's Guide Through Latent Space
by: von Rütte, Dimitri, et al.
Published: (2024)
by: von Rütte, Dimitri, et al.
Published: (2024)
Light Unbalanced Optimal Transport
by: Gazdieva, Milena, et al.
Published: (2023)
by: Gazdieva, Milena, et al.
Published: (2023)
Sliced-Regularized Optimal Transport
by: Nguyen, Khai
Published: (2026)
by: Nguyen, Khai
Published: (2026)
Variational Entropic Optimal Transport
by: Dyachenko, Roman, et al.
Published: (2026)
by: Dyachenko, Roman, et al.
Published: (2026)
Simplifying Transformer Blocks
by: He, Bobby, et al.
Published: (2023)
by: He, Bobby, et al.
Published: (2023)
A Specialized Semismooth Newton Method for Kernel-Based Optimal Transport
by: Lin, Tianyi, et al.
Published: (2023)
by: Lin, Tianyi, et al.
Published: (2023)
Optimal Transport under Group Fairness Constraints
by: Bleistein, Linus, et al.
Published: (2026)
by: Bleistein, Linus, et al.
Published: (2026)
Similar Items
-
Towards Meta-Pruning via Optimal Transport
by: Theus, Alexander, et al.
Published: (2024) -
Model Fusion via Retrofitting
by: Luenam, Phoomraphee, et al.
Published: (2025) -
Generalized Linear Mode Connectivity for Transformers
by: Theus, Alexander, et al.
Published: (2025) -
Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
by: Anagnostidis, Sotiris, et al.
Published: (2023) -
How Susceptible are LLMs to Influence in Prompts?
by: Anagnostidis, Sotiris, et al.
Published: (2024)