The Non-Local Model Merging Problem: Permutation Symmetries and Variance Collapse
Fuente:
arXiv
Saved in:
| Main Authors: | Sharma, Ekansh, Roy, Daniel M., Dziugaite, Gintare Karolina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simultaneous linear connectivity of neural networks modulo permutation
by: Sharma, Ekansh, et al.
Published: (2024)
by: Sharma, Ekansh, et al.
Published: (2024)
Information Complexity of Stochastic Convex Optimization: Applications to Generalization and Memorization
by: Attias, Idan, et al.
Published: (2024)
by: Attias, Idan, et al.
Published: (2024)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Improved Localized Machine Unlearning Through the Lens of Memorization
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024)
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024)
Continual Learning in Vision-Language Models via Aligned Model Merging
by: Sokar, Ghada, et al.
Published: (2025)
by: Sokar, Ghada, et al.
Published: (2025)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
by: Guo, Phillip, et al.
Published: (2024)
by: Guo, Phillip, et al.
Published: (2024)
On Traceability in $\ell_p$ Stochastic Convex Optimization
by: Voitovych, Sasha, et al.
Published: (2025)
by: Voitovych, Sasha, et al.
Published: (2025)
Unlearning in- vs. out-of-distribution data in LLMs under gradient-based method
by: Baluta, Teodora, et al.
Published: (2024)
by: Baluta, Teodora, et al.
Published: (2024)
Data Selection for Transfer Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024)
Detoxifying LLMs via Representation Erasure-Based Preference Optimization
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2026)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2026)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
by: Yang, Yu, et al.
Published: (2023)
by: Yang, Yu, et al.
Published: (2023)
Dataset Difficulty and the Role of Inductive Bias
by: Kwok, Devin, et al.
Published: (2024)
by: Kwok, Devin, et al.
Published: (2024)
Leveraging Function Space Aggregation for Federated Learning at Scale
by: Dhawan, Nikita, et al.
Published: (2023)
by: Dhawan, Nikita, et al.
Published: (2023)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
Soup to go: mitigating forgetting during continual learning with model averaging
by: Kleiman, Anat, et al.
Published: (2025)
by: Kleiman, Anat, et al.
Published: (2025)
Leveraging Per-Instance Privacy for Machine Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
SSFL: Discovering Sparse Unified Subnetworks at Initialization for Efficient Federated Learning
by: Ohib, Riyasat, et al.
Published: (2024)
by: Ohib, Riyasat, et al.
Published: (2024)
Torque-Aware Momentum
by: Malviya, Pranshu, et al.
Published: (2024)
by: Malviya, Pranshu, et al.
Published: (2024)
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
by: Jin, Tian, et al.
Published: (2025)
by: Jin, Tian, et al.
Published: (2025)
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
Sparse Training from Random Initialization: Aligning Lottery Ticket Masks using Weight Symmetry
by: Adnan, Mohammed, et al.
Published: (2025)
by: Adnan, Mohammed, et al.
Published: (2025)
PLeaS -- Merging Models with Permutations and Least Squares
by: Nasery, Anshul, et al.
Published: (2024)
by: Nasery, Anshul, et al.
Published: (2024)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
by: Inane, Ahmed Mehdi, et al.
Published: (2026)
by: Inane, Ahmed Mehdi, et al.
Published: (2026)
Merge Now, Regret Later: The Hidden Cost of Model Merging is Adversarial Transferability
by: Gangwal, Ankit, et al.
Published: (2025)
by: Gangwal, Ankit, et al.
Published: (2025)
Mixture of Experts in a Mixture of RL settings
by: Willi, Timon, et al.
Published: (2024)
by: Willi, Timon, et al.
Published: (2024)
Zeroth-Order Non-Log-Concave Sampling with Variance Reduction and Applications to Inverse Problems
by: Sahin, M. Berk, et al.
Published: (2026)
by: Sahin, M. Berk, et al.
Published: (2026)
Beyond the Permutation Symmetry of Transformers: The Role of Rotation for Model Fusion
by: Zhang, Binchi, et al.
Published: (2025)
by: Zhang, Binchi, et al.
Published: (2025)
Variational Inference Failures Under Model Symmetries: Permutation Invariant Posteriors for Bayesian Neural Networks
by: Gelberg, Yoav, et al.
Published: (2024)
by: Gelberg, Yoav, et al.
Published: (2024)
Lost in Translation: How Language Re-Aligns Vision for Cross-Species Pathology
by: Arora, Ekansh
Published: (2026)
by: Arora, Ekansh
Published: (2026)
Are we making progress in unlearning? Findings from the first NeurIPS unlearning competition
by: Triantafillou, Eleni, et al.
Published: (2024)
by: Triantafillou, Eleni, et al.
Published: (2024)
VARSHAP: Addressing Global Dependency Problems in Explainable AI with Variance-Based Local Feature Attribution
by: Gajewski, Mateusz, et al.
Published: (2025)
by: Gajewski, Mateusz, et al.
Published: (2025)
On the Extreme Variance of Certified Local Robustness Across Model Seeds
by: Le, Minh, et al.
Published: (2026)
by: Le, Minh, et al.
Published: (2026)
Permutation Picture of Graph Combinatorial Optimization Problems
by: Min, Yimeng
Published: (2024)
by: Min, Yimeng
Published: (2024)
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
A Compact Representation for Bayesian Neural Networks By Removing Permutation Symmetry
by: Xiao, Tim Z., et al.
Published: (2023)
by: Xiao, Tim Z., et al.
Published: (2023)
Symmetry-Aware Graph Metanetwork Autoencoders: Model Merging through Parameter Canonicalization
by: Boufalis, Odysseas, et al.
Published: (2025)
by: Boufalis, Odysseas, et al.
Published: (2025)
Rank, Head-Channel Non-Identifiability, and Symmetry Breaking: A Precise Analysis of Representational Collapse in Transformers
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
CavMerge: Merging K-means Based on Local Log-Concavity
by: Qiao, Zhili, et al.
Published: (2026)
by: Qiao, Zhili, et al.
Published: (2026)
Communication-Efficient Personalized Adaptation via Federated-Local Model Merging
by: Zou, Yinan, et al.
Published: (2026)
by: Zou, Yinan, et al.
Published: (2026)
Similar Items
-
Simultaneous linear connectivity of neural networks modulo permutation
by: Sharma, Ekansh, et al.
Published: (2024) -
Information Complexity of Stochastic Convex Optimization: Applications to Generalization and Memorization
by: Attias, Idan, et al.
Published: (2024) -
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025) -
Improved Localized Machine Unlearning Through the Lens of Memorization
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024) -
Continual Learning in Vision-Language Models via Aligned Model Merging
by: Sokar, Ghada, et al.
Published: (2025)