How does the optimizer implicitly bias the model merging loss landscape?
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chenxiang, Theus, Alexander, Teney, Damien, Orvieto, Antonio, Pang, Jun, Mauw, Sjouke |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spurious Privacy Leakage in Neural Networks
by: Zhang, Chenxiang, et al.
Published: (2025)
by: Zhang, Chenxiang, et al.
Published: (2025)
Bits for Privacy: Evaluating Post-Training Quantization via Membership Inference
by: Zhang, Chenxiang, et al.
Published: (2025)
by: Zhang, Chenxiang, et al.
Published: (2025)
Meta-RL Induces Exploration in Language Agents
by: Jiang, Yulun, et al.
Published: (2025)
by: Jiang, Yulun, et al.
Published: (2025)
Scalable Ensemble Diversification for OOD Generalization and Detection
by: Rubinstein, Alexander, et al.
Published: (2024)
by: Rubinstein, Alexander, et al.
Published: (2024)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
by: Zucchet, Nicolas, et al.
Published: (2024)
by: Zucchet, Nicolas, et al.
Published: (2024)
Neural Redshift: Random Networks are not Random Functions
by: Teney, Damien, et al.
Published: (2024)
by: Teney, Damien, et al.
Published: (2024)
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
by: Scimeca, Luca, et al.
Published: (2023)
by: Scimeca, Luca, et al.
Published: (2023)
Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss
by: de Frutos, José Manuel, et al.
Published: (2024)
by: de Frutos, José Manuel, et al.
Published: (2024)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
by: Elhassan, Fay, et al.
Published: (2025)
by: Elhassan, Fay, et al.
Published: (2025)
Design Principles for Sequence Models via Coefficient Dynamics
by: Sieber, Jerome, et al.
Published: (2025)
by: Sieber, Jerome, et al.
Published: (2025)
Recovering implicit physics model under real-world constraints
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
How does Bayesian Sampling help Membership Inference Attacks?
by: Liu, Zhenlong, et al.
Published: (2025)
by: Liu, Zhenlong, et al.
Published: (2025)
Causal vs. Anticausal merging of predictors
by: Mejia, Sergio Hernan Garrido, et al.
Published: (2025)
by: Mejia, Sergio Hernan Garrido, et al.
Published: (2025)
Generalized Linear Mode Connectivity for Transformers
by: Theus, Alexander, et al.
Published: (2025)
by: Theus, Alexander, et al.
Published: (2025)
What model does MuZero learn?
by: He, Jinke, et al.
Published: (2023)
by: He, Jinke, et al.
Published: (2023)
Rolling Ball Optimizer: Learning by ironing out loss landscape wrinkles
by: Belgoumri, Mohammed Djameleddine, et al.
Published: (2025)
by: Belgoumri, Mohammed Djameleddine, et al.
Published: (2025)
OOD-Chameleon: Is Algorithm Selection for OOD Generalization Learnable?
by: Jiang, Liangze, et al.
Published: (2024)
by: Jiang, Liangze, et al.
Published: (2024)
Fréchet regression with implicit denoising and multicollinearity reduction
by: Mansouri, Dou El Kefel, et al.
Published: (2024)
by: Mansouri, Dou El Kefel, et al.
Published: (2024)
Closed-form merging of parameter-efficient modules for Federated Continual Learning
by: Salami, Riccardo, et al.
Published: (2024)
by: Salami, Riccardo, et al.
Published: (2024)
(Almost) Free Modality Stitching of Foundation Models
by: Singh, Jaisidh, et al.
Published: (2025)
by: Singh, Jaisidh, et al.
Published: (2025)
How to systematically develop an effective AI-based bias correction model?
by: Zhou, Xiao, et al.
Published: (2025)
by: Zhou, Xiao, et al.
Published: (2025)
Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone Ensembling
by: Rodriguez-Opazo, Cristian, et al.
Published: (2024)
by: Rodriguez-Opazo, Cristian, et al.
Published: (2024)
Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
by: Sieber, Jerome, et al.
Published: (2024)
by: Sieber, Jerome, et al.
Published: (2024)
How to be fair? A study of label and selection bias
by: Favier, Marco, et al.
Published: (2024)
by: Favier, Marco, et al.
Published: (2024)
Uncertainty modeling for fine-tuned implicit functions
by: Susmelj, Anna, et al.
Published: (2024)
by: Susmelj, Anna, et al.
Published: (2024)
Enhancing robustness of data-driven SHM models: adversarial training with circle loss
by: Yang, Xiangli, et al.
Published: (2024)
by: Yang, Xiangli, et al.
Published: (2024)
Generalized Interpolating Discrete Diffusion
by: von Rütte, Dimitri, et al.
Published: (2025)
by: von Rütte, Dimitri, et al.
Published: (2025)
GDP nowcasting with artificial neural networks: How much does long-term memory matter?
by: Németh, Kristóf, et al.
Published: (2023)
by: Németh, Kristóf, et al.
Published: (2023)
Exploring multimodal implicit behavior learning for vehicle navigation in simulated cities
by: Antonelo, Eric Aislan, et al.
Published: (2025)
by: Antonelo, Eric Aislan, et al.
Published: (2025)
GRASP: Deterministic argument ranking in interaction graphs
by: Misra, Diganta, et al.
Published: (2026)
by: Misra, Diganta, et al.
Published: (2026)
Understanding the dynamics of the frequency bias in neural networks
by: Molina, Juan, et al.
Published: (2024)
by: Molina, Juan, et al.
Published: (2024)
Detecting labeling bias using influence functions
by: Jørgensen, Frida, et al.
Published: (2026)
by: Jørgensen, Frida, et al.
Published: (2026)
Improving the classification of extreme classes by means of loss regularisation and generalised beta distributions
by: Vargas, Víctor Manuel, et al.
Published: (2024)
by: Vargas, Víctor Manuel, et al.
Published: (2024)
On student-teacher deviations in distillation: does it pay to disobey?
by: Nagarajan, Vaishnavh, et al.
Published: (2023)
by: Nagarajan, Vaishnavh, et al.
Published: (2023)
Chemist-aligned retrosynthesis by ensembling diverse inductive bias models
by: Maziarz, Krzysztof, et al.
Published: (2024)
by: Maziarz, Krzysztof, et al.
Published: (2024)
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023)
by: Coda-Forno, Julian, et al.
Published: (2023)
Text-guided multi-property molecular optimization with a diffusion language model
by: Xiong, Yida, et al.
Published: (2024)
by: Xiong, Yida, et al.
Published: (2024)
Generative method for aerodynamic optimization based on classifier-free guided denoising diffusion probabilistic model
by: Deng, Shisong, et al.
Published: (2025)
by: Deng, Shisong, et al.
Published: (2025)
Bayesian Low-Rank LeArning (Bella): A Practical Approach to Bayesian Neural Networks
by: Doan, Bao Gia, et al.
Published: (2024)
by: Doan, Bao Gia, et al.
Published: (2024)
Towards certifiable AI in aviation: landscape, challenges, and opportunities
by: Bello, Hymalai, et al.
Published: (2024)
by: Bello, Hymalai, et al.
Published: (2024)
Similar Items
-
Spurious Privacy Leakage in Neural Networks
by: Zhang, Chenxiang, et al.
Published: (2025) -
Bits for Privacy: Evaluating Post-Training Quantization via Membership Inference
by: Zhang, Chenxiang, et al.
Published: (2025) -
Meta-RL Induces Exploration in Language Agents
by: Jiang, Yulun, et al.
Published: (2025) -
Scalable Ensemble Diversification for OOD Generalization and Detection
by: Rubinstein, Alexander, et al.
Published: (2024) -
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
by: Zucchet, Nicolas, et al.
Published: (2024)