Generalized Linear Mode Connectivity for Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Theus, Alexander, Cabodi, Alessandro, Anagnostidis, Sotiris, Orvieto, Antonio, Singh, Sidak Pal, Boeva, Valentina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Meta-Pruning via Optimal Transport
by: Theus, Alexander, et al.
Published: (2024)
by: Theus, Alexander, et al.
Published: (2024)
Transformer Fusion with Optimal Transport
by: Imfeld, Moritz, et al.
Published: (2023)
by: Imfeld, Moritz, et al.
Published: (2023)
Landscaping Linear Mode Connectivity
by: Singh, Sidak Pal, et al.
Published: (2024)
by: Singh, Sidak Pal, et al.
Published: (2024)
Model Fusion via Retrofitting
by: Luenam, Phoomraphee, et al.
Published: (2025)
by: Luenam, Phoomraphee, et al.
Published: (2025)
How Susceptible are LLMs to Influence in Prompts?
by: Anagnostidis, Sotiris, et al.
Published: (2024)
by: Anagnostidis, Sotiris, et al.
Published: (2024)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
by: Ormaniec, Weronika, et al.
Published: (2024)
by: Ormaniec, Weronika, et al.
Published: (2024)
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025)
by: Zakarin, Daniyar, et al.
Published: (2025)
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
by: Khromov, Grigory, et al.
Published: (2023)
by: Khromov, Grigory, et al.
Published: (2023)
How does the optimizer implicitly bias the model merging loss landscape?
by: Zhang, Chenxiang, et al.
Published: (2025)
by: Zhang, Chenxiang, et al.
Published: (2025)
Explaining Grokking in Transformers through the Lens of Inductive Bias
by: Singh, Jaisidh, et al.
Published: (2026)
by: Singh, Jaisidh, et al.
Published: (2026)
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
by: Anagnostidis, Sotiris, et al.
Published: (2023)
by: Anagnostidis, Sotiris, et al.
Published: (2023)
Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers
by: Bill, Eric Tillman, et al.
Published: (2025)
by: Bill, Eric Tillman, et al.
Published: (2025)
Feature Clock: High-Dimensional Effects in Two-Dimensional Plots
by: Ovcharenko, Olga, et al.
Published: (2024)
by: Ovcharenko, Olga, et al.
Published: (2024)
Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
by: Anagnostidis, Sotiris, et al.
Published: (2023)
by: Anagnostidis, Sotiris, et al.
Published: (2023)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
An Uncertainty Principle for Linear Recurrent Neural Networks
by: François, Alexandre, et al.
Published: (2025)
by: François, Alexandre, et al.
Published: (2025)
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025)
by: Singh, Sidak Pal, et al.
Published: (2025)
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
by: Singh, Sidak Pal, et al.
Published: (2024)
by: Singh, Sidak Pal, et al.
Published: (2024)
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
by: Belloni, Annalisa, et al.
Published: (2026)
by: Belloni, Annalisa, et al.
Published: (2026)
Local vs Global continual learning
by: Lanzillotta, Giulia, et al.
Published: (2024)
by: Lanzillotta, Giulia, et al.
Published: (2024)
Adam Simplified: Bias Correction Debunked
by: Laing, Sam, et al.
Published: (2025)
by: Laing, Sam, et al.
Published: (2025)
Revisiting associative recall in modern recurrent models
by: Okpekpe, Destiny, et al.
Published: (2025)
by: Okpekpe, Destiny, et al.
Published: (2025)
Layer-wise Linear Mode Connectivity
by: Adilova, Linara, et al.
Published: (2023)
by: Adilova, Linara, et al.
Published: (2023)
FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute
by: Anagnostidis, Sotiris, et al.
Published: (2025)
by: Anagnostidis, Sotiris, et al.
Published: (2025)
On Linear Mode Connectivity of Mixture-of-Experts Architectures
by: Tran, Viet-Hoang, et al.
Published: (2025)
by: Tran, Viet-Hoang, et al.
Published: (2025)
Linear Mode Connectivity in Differentiable Tree Ensembles
by: Kanoh, Ryuichi, et al.
Published: (2024)
by: Kanoh, Ryuichi, et al.
Published: (2024)
In Search of Adam's Secret Sauce
by: Orvieto, Antonio, et al.
Published: (2025)
by: Orvieto, Antonio, et al.
Published: (2025)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
by: Orvieto, Antonio, et al.
Published: (2024)
by: Orvieto, Antonio, et al.
Published: (2024)
scTree: Discovering Cellular Hierarchies in the Presence of Batch Effects in scRNA-seq Data
by: Vandenhirtz, Moritz, et al.
Published: (2024)
by: Vandenhirtz, Moritz, et al.
Published: (2024)
Analyzing the Role of Permutation Invariance in Linear Mode Connectivity
by: Zhan, Keyao, et al.
Published: (2025)
by: Zhan, Keyao, et al.
Published: (2025)
Linear Mode Connectivity in Sparse Neural Networks
by: McDermott, Luke, et al.
Published: (2023)
by: McDermott, Luke, et al.
Published: (2023)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
by: Zucchet, Nicolas, et al.
Published: (2024)
by: Zucchet, Nicolas, et al.
Published: (2024)
When, Where and Why to Average Weights?
by: Ajroldi, Niccolò, et al.
Published: (2025)
by: Ajroldi, Niccolò, et al.
Published: (2025)
Improved state mixing in higher-order and block diagonal linear recurrent networks
by: Dubinin, Igor, et al.
Published: (2026)
by: Dubinin, Igor, et al.
Published: (2026)
Proving Linear Mode Connectivity of Neural Networks via Optimal Transport
by: Ferbach, Damien, et al.
Published: (2023)
by: Ferbach, Damien, et al.
Published: (2023)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
by: Srećković, Teodora, et al.
Published: (2025)
by: Srećković, Teodora, et al.
Published: (2025)
Label Attention Network for Temporal Sets Prediction: You Were Looking at a Wrong Self-Attention
by: Kovtun, Elizaveta, et al.
Published: (2023)
by: Kovtun, Elizaveta, et al.
Published: (2023)
Linear Mode Connectivity under Data Shifts for Deep Ensembles of Image Classifiers
by: Hepburn, C., et al.
Published: (2025)
by: Hepburn, C., et al.
Published: (2025)
(Almost) Free Modality Stitching of Foundation Models
by: Singh, Jaisidh, et al.
Published: (2025)
by: Singh, Jaisidh, et al.
Published: (2025)
Similar Items
-
Towards Meta-Pruning via Optimal Transport
by: Theus, Alexander, et al.
Published: (2024) -
Transformer Fusion with Optimal Transport
by: Imfeld, Moritz, et al.
Published: (2023) -
Landscaping Linear Mode Connectivity
by: Singh, Sidak Pal, et al.
Published: (2024) -
Model Fusion via Retrofitting
by: Luenam, Phoomraphee, et al.
Published: (2025) -
How Susceptible are LLMs to Influence in Prompts?
by: Anagnostidis, Sotiris, et al.
Published: (2024)