Landscaping Linear Mode Connectivity
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Sidak Pal, Adilova, Linara, Kamp, Michael, Fischer, Asja, Schölkopf, Bernhard, Hofmann, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Layer-wise Linear Mode Connectivity
by: Adilova, Linara, et al.
Published: (2023)
by: Adilova, Linara, et al.
Published: (2023)
The Uncanny Valley: Exploring Adversarial Robustness from a Flatness Perspective
by: Walter, Nils Philipp, et al.
Published: (2024)
by: Walter, Nils Philipp, et al.
Published: (2024)
When Flatness Does (Not) Guarantee Adversarial Robustness
by: Walter, Nils Philipp, et al.
Published: (2025)
by: Walter, Nils Philipp, et al.
Published: (2025)
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
by: Singh, Sidak Pal, et al.
Published: (2024)
by: Singh, Sidak Pal, et al.
Published: (2024)
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
by: Han, Ting, et al.
Published: (2025)
by: Han, Ting, et al.
Published: (2025)
Generalized Linear Mode Connectivity for Transformers
by: Theus, Alexander, et al.
Published: (2025)
by: Theus, Alexander, et al.
Published: (2025)
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
by: Khromov, Grigory, et al.
Published: (2023)
by: Khromov, Grigory, et al.
Published: (2023)
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025)
by: Zakarin, Daniyar, et al.
Published: (2025)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
by: Ormaniec, Weronika, et al.
Published: (2024)
by: Ormaniec, Weronika, et al.
Published: (2024)
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Local vs Global continual learning
by: Lanzillotta, Giulia, et al.
Published: (2024)
by: Lanzillotta, Giulia, et al.
Published: (2024)
Transformer Fusion with Optimal Transport
by: Imfeld, Moritz, et al.
Published: (2023)
by: Imfeld, Moritz, et al.
Published: (2023)
Fisher information flow in artificial neural networks
by: Weimar, Maximilian, et al.
Published: (2025)
by: Weimar, Maximilian, et al.
Published: (2025)
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025)
by: Singh, Sidak Pal, et al.
Published: (2025)
Towards Meta-Pruning via Optimal Transport
by: Theus, Alexander, et al.
Published: (2024)
by: Theus, Alexander, et al.
Published: (2024)
Geometry-Aware Instrumental Variable Regression
by: Kremer, Heiner, et al.
Published: (2024)
by: Kremer, Heiner, et al.
Published: (2024)
Robustness of Nonlinear Representation Learning
by: Buchholz, Simon, et al.
Published: (2025)
by: Buchholz, Simon, et al.
Published: (2025)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
On the regularization of Wasserstein GANs
by: Petzka, Henning, et al.
Published: (2017)
by: Petzka, Henning, et al.
Published: (2017)
Model Fusion via Retrofitting
by: Luenam, Phoomraphee, et al.
Published: (2025)
by: Luenam, Phoomraphee, et al.
Published: (2025)
Linear Mode Connectivity in Differentiable Tree Ensembles
by: Kanoh, Ryuichi, et al.
Published: (2024)
by: Kanoh, Ryuichi, et al.
Published: (2024)
On Linear Mode Connectivity of Mixture-of-Experts Architectures
by: Tran, Viet-Hoang, et al.
Published: (2025)
by: Tran, Viet-Hoang, et al.
Published: (2025)
Skill or Luck? Return Decomposition via Advantage Functions
by: Pan, Hsiao-Ru, et al.
Published: (2024)
by: Pan, Hsiao-Ru, et al.
Published: (2024)
PENEX: AdaBoost-Inspired Neural Network Regularization
by: Kladny, Klaus-Rudolf, et al.
Published: (2025)
by: Kladny, Klaus-Rudolf, et al.
Published: (2025)
Analyzing the Role of Permutation Invariance in Linear Mode Connectivity
by: Zhan, Keyao, et al.
Published: (2025)
by: Zhan, Keyao, et al.
Published: (2025)
Physics of Learning: A Lagrangian perspective to different learning paradigms
by: Guo, Siyuan, et al.
Published: (2025)
by: Guo, Siyuan, et al.
Published: (2025)
Linear Mode Connectivity in Sparse Neural Networks
by: McDermott, Luke, et al.
Published: (2023)
by: McDermott, Luke, et al.
Published: (2023)
Learning Sparse Codes with Entropy-Based ELBOs
by: Velychko, Dmytro, et al.
Published: (2023)
by: Velychko, Dmytro, et al.
Published: (2023)
SPARTAN: A Sparse Transformer World Model Attending to What Matters
by: Lei, Anson, et al.
Published: (2024)
by: Lei, Anson, et al.
Published: (2024)
Causal Modeling with Stationary Diffusions
by: Lorch, Lars, et al.
Published: (2023)
by: Lorch, Lars, et al.
Published: (2023)
Online Learning and Unlearning
by: Hu, Yaxi, et al.
Published: (2025)
by: Hu, Yaxi, et al.
Published: (2025)
Targeted Reduction of Causal Models
by: Kekić, Armin, et al.
Published: (2023)
by: Kekić, Armin, et al.
Published: (2023)
Conformal Generative Modeling with Improved Sample Efficiency through Sequential Greedy Filtering
by: Kladny, Klaus-Rudolf, et al.
Published: (2024)
by: Kladny, Klaus-Rudolf, et al.
Published: (2024)
Comparative Study on Noise-Augmented Training and its Effect on Adversarial Robustness in ASR Systems
by: Pizzi, Karla, et al.
Published: (2024)
by: Pizzi, Karla, et al.
Published: (2024)
Robustifying automatic speech recognition by extracting slowly varying features
by: Pizarro, Matías, et al.
Published: (2021)
by: Pizarro, Matías, et al.
Published: (2021)
Precision-Varying Prediction (PVP): Robustifying ASR systems against adversarial attacks
by: Pizarro, Matías, et al.
Published: (2026)
by: Pizarro, Matías, et al.
Published: (2026)
Adaptive Inverted-Index Routing for Granular Mixtures-of-Experts
by: Kladny, Klaus-Rudolf, et al.
Published: (2026)
by: Kladny, Klaus-Rudolf, et al.
Published: (2026)
Proving Linear Mode Connectivity of Neural Networks via Optimal Transport
by: Ferbach, Damien, et al.
Published: (2023)
by: Ferbach, Damien, et al.
Published: (2023)
Scaling Behavior of Discrete Diffusion Language Models
by: von Rütte, Dimitri, et al.
Published: (2025)
by: von Rütte, Dimitri, et al.
Published: (2025)
DistriBlock: Identifying adversarial audio samples by leveraging characteristics of the output distribution
by: Pizarro, Matías, et al.
Published: (2023)
by: Pizarro, Matías, et al.
Published: (2023)
Similar Items
-
Layer-wise Linear Mode Connectivity
by: Adilova, Linara, et al.
Published: (2023) -
The Uncanny Valley: Exploring Adversarial Robustness from a Flatness Perspective
by: Walter, Nils Philipp, et al.
Published: (2024) -
When Flatness Does (Not) Guarantee Adversarial Robustness
by: Walter, Nils Philipp, et al.
Published: (2025) -
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
by: Singh, Sidak Pal, et al.
Published: (2024) -
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
by: Han, Ting, et al.
Published: (2025)