Model Merging by Uncertainty-Based Gradient Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Daheim, Nico, Möllenhoff, Thomas, Ponti, Edoardo Maria, Gurevych, Iryna, Khan, Mohammad Emtiyaz |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncertainty-Aware Decoding with Minimum Bayes Risk
by: Daheim, Nico, et al.
Published: (2025)
by: Daheim, Nico, et al.
Published: (2025)
How to Weight Multitask Finetuning? Fast Previews via Bayesian Model-Merging
by: Maldonado, Hugo Monzón, et al.
Published: (2024)
by: Maldonado, Hugo Monzón, et al.
Published: (2024)
Improving LoRA with Variational Learning
by: Cong, Bai, et al.
Published: (2025)
by: Cong, Bai, et al.
Published: (2025)
Variational Low-Rank Adaptation Using IVON
by: Cong, Bai, et al.
Published: (2024)
by: Cong, Bai, et al.
Published: (2024)
SVRG and Beyond via Posterior Correction
by: Daheim, Nico, et al.
Published: (2025)
by: Daheim, Nico, et al.
Published: (2025)
Variational Learning is Effective for Large Deep Networks
by: Shen, Yuesong, et al.
Published: (2024)
by: Shen, Yuesong, et al.
Published: (2024)
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
by: Daheim, Nico, et al.
Published: (2024)
by: Daheim, Nico, et al.
Published: (2024)
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
by: Macina, Jakub, et al.
Published: (2025)
by: Macina, Jakub, et al.
Published: (2025)
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
by: Kumar, Navish, et al.
Published: (2025)
by: Kumar, Navish, et al.
Published: (2025)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
by: Tamoyan, Hovhannes, et al.
Published: (2025)
by: Tamoyan, Hovhannes, et al.
Published: (2025)
The Memory Perturbation Equation: Understanding Model's Sensitivity to Data
by: Nickl, Peter, et al.
Published: (2023)
by: Nickl, Peter, et al.
Published: (2023)
SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modelling
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2025)
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2025)
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2024)
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2024)
From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning
by: Dinucu-Jianu, David, et al.
Published: (2025)
by: Dinucu-Jianu, David, et al.
Published: (2025)
Token Weighting for Long-Range Language Modeling
by: Helm, Falko, et al.
Published: (2025)
by: Helm, Falko, et al.
Published: (2025)
Probing the Emergence of Cross-lingual Alignment during LLM Training
by: Wang, Hetong, et al.
Published: (2024)
by: Wang, Hetong, et al.
Published: (2024)
Towards Automated Error Discovery: A Study in Conversational AI
by: Petrak, Dominic, et al.
Published: (2025)
by: Petrak, Dominic, et al.
Published: (2025)
Knowledge Adaptation as Posterior Correction
by: Khan, Mohammad Emtiyaz
Published: (2025)
by: Khan, Mohammad Emtiyaz
Published: (2025)
Information Geometry of Variational Bayes
by: Khan, Mohammad Emtiyaz
Published: (2025)
by: Khan, Mohammad Emtiyaz
Published: (2025)
Scaling Sparse Fine-Tuning to Large Language Models
by: Ansell, Alan, et al.
Published: (2024)
by: Ansell, Alan, et al.
Published: (2024)
Emergent Communication Pretraining for Few-Shot Machine Translation
by: Li, Yaoyiran, et al.
Published: (2020)
by: Li, Yaoyiran, et al.
Published: (2020)
Spectral Editing of Activations for Large Language Model Alignment
by: Qiu, Yifu, et al.
Published: (2024)
by: Qiu, Yifu, et al.
Published: (2024)
Mixtures of In-Context Learners
by: Hong, Giwon, et al.
Published: (2024)
by: Hong, Giwon, et al.
Published: (2024)
Self-Improving World Modelling with Latent Actions
by: Qiu, Yifu, et al.
Published: (2026)
by: Qiu, Yifu, et al.
Published: (2026)
Merge to Mix: Mixing Datasets via Model Merging
by: Tao, Zhixu Silvia, et al.
Published: (2025)
by: Tao, Zhixu Silvia, et al.
Published: (2025)
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
by: Lu, Zhenyi, et al.
Published: (2024)
by: Lu, Zhenyi, et al.
Published: (2024)
Arcee's MergeKit: A Toolkit for Merging Large Language Models
by: Goddard, Charles, et al.
Published: (2024)
by: Goddard, Charles, et al.
Published: (2024)
Socratic Reasoning Improves Positive Text Rewriting
by: Goel, Anmol, et al.
Published: (2024)
by: Goel, Anmol, et al.
Published: (2024)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
by: Huang, Zeyu, et al.
Published: (2025)
by: Huang, Zeyu, et al.
Published: (2025)
Model Merging for Knowledge Editing
by: Fu, Zichuan, et al.
Published: (2025)
by: Fu, Zichuan, et al.
Published: (2025)
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
by: Shenaj, Donald, et al.
Published: (2025)
by: Shenaj, Donald, et al.
Published: (2025)
Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
by: Yang, Jinluan, et al.
Published: (2025)
by: Yang, Jinluan, et al.
Published: (2025)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
by: Baumgärtner, Tim, et al.
Published: (2026)
by: Baumgärtner, Tim, et al.
Published: (2026)
An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms
by: Grünefeld, Nils, et al.
Published: (2026)
by: Grünefeld, Nils, et al.
Published: (2026)
LLM-Match: An Open-Sourced Patient Matching Model Based on Large Language Models and Retrieval-Augmented Generation
by: Li, Xiaodi, et al.
Published: (2025)
by: Li, Xiaodi, et al.
Published: (2025)
Merging in a Bottle: Differentiable Adaptive Merging (DAM) and the Path from Averaging to Automation
by: Gauthier-Caron, Thomas, et al.
Published: (2024)
by: Gauthier-Caron, Thomas, et al.
Published: (2024)
Dynamic Model Merging Made Slim
by: Du, Guodong, et al.
Published: (2026)
by: Du, Guodong, et al.
Published: (2026)
What Matters for Model Merging at Scale?
by: Yadav, Prateek, et al.
Published: (2024)
by: Yadav, Prateek, et al.
Published: (2024)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
by: Lai, Kunfeng, et al.
Published: (2025)
by: Lai, Kunfeng, et al.
Published: (2025)
Channel Merging: Preserving Specialization for Merged Experts
by: Zhang, Mingyang, et al.
Published: (2024)
by: Zhang, Mingyang, et al.
Published: (2024)
Similar Items
-
Uncertainty-Aware Decoding with Minimum Bayes Risk
by: Daheim, Nico, et al.
Published: (2025) -
How to Weight Multitask Finetuning? Fast Previews via Bayesian Model-Merging
by: Maldonado, Hugo Monzón, et al.
Published: (2024) -
Improving LoRA with Variational Learning
by: Cong, Bai, et al.
Published: (2025) -
Variational Low-Rank Adaptation Using IVON
by: Cong, Bai, et al.
Published: (2024) -
SVRG and Beyond via Posterior Correction
by: Daheim, Nico, et al.
Published: (2025)