Model Merging via Data-Free Covariance Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Hameed, Marawan Gamal Abdel, Tam, Derek, Notsawo, Pascal Jr Tikeng, Raffel, Colin, Rabusseau, Guillaume |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grokking Beyond the Euclidean Norm of Model Parameters
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2025)
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2025)
Grokking Finite-Dimensional Algebra
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2026)
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2026)
Efficient Probabilistic Tensor Networks
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2025)
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2025)
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2024)
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2024)
A Tensor Decomposition Perspective on Second-order RNNs
by: Lizaire, Maude, et al.
Published: (2024)
by: Lizaire, Maude, et al.
Published: (2024)
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023)
by: Tam, Derek, et al.
Published: (2023)
Realistic Evaluation of Model Merging for Compositional Generalization
by: Tam, Derek, et al.
Published: (2024)
by: Tam, Derek, et al.
Published: (2024)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
by: Li, YuXin, et al.
Published: (2025)
by: Li, YuXin, et al.
Published: (2025)
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
by: Je, Gyung Hyun, et al.
Published: (2025)
by: Je, Gyung Hyun, et al.
Published: (2025)
Soft Merging of Experts with Adaptive Routing
by: Muqeeth, Mohammed, et al.
Published: (2023)
by: Muqeeth, Mohammed, et al.
Published: (2023)
Lost in Translation: The Algorithmic Gap Between LMs and the Brain
by: Tosato, Tommaso, et al.
Published: (2024)
by: Tosato, Tommaso, et al.
Published: (2024)
Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
by: Matena, Michael, et al.
Published: (2023)
by: Matena, Michael, et al.
Published: (2023)
Tractable Shapley Values and Interactions via Tensor Networks
by: Heidari, Farzaneh, et al.
Published: (2025)
by: Heidari, Farzaneh, et al.
Published: (2025)
Tensor Cookbook: Mastering Tensors through Diagrams
by: Rakhshan, Beheshteh T., et al.
Published: (2026)
by: Rakhshan, Beheshteh T., et al.
Published: (2026)
TN-SHAP-G: Graph-Structured Tensor Network Surrogates for Shapley Values and Interactions
by: Heidari, Farzaneh, et al.
Published: (2026)
by: Heidari, Farzaneh, et al.
Published: (2026)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Higher Order Transformers: Enhancing Stock Movement Prediction On Multimodal Time-Series Data
by: Omranpour, Soroush, et al.
Published: (2024)
by: Omranpour, Soroush, et al.
Published: (2024)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
by: Lesens, Damien, et al.
Published: (2025)
by: Lesens, Damien, et al.
Published: (2025)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
by: Rizvi-Martel, Michael, et al.
Published: (2026)
by: Rizvi-Martel, Michael, et al.
Published: (2026)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
by: Liu, Haokun, et al.
Published: (2026)
by: Liu, Haokun, et al.
Published: (2026)
Higher-Order Transformers With Kronecker-Structured Attention
by: Omranpour, Soroush, et al.
Published: (2024)
by: Omranpour, Soroush, et al.
Published: (2024)
Enhancing Training Data Attribution with Representational Optimization
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
FlowQ-Net: A Generative Framework for Automated Quantum Circuit Design
by: Dai, Jun, et al.
Published: (2025)
by: Dai, Jun, et al.
Published: (2025)
Learning to Route Among Specialized Experts for Zero-Shot Generalization
by: Muqeeth, Mohammed, et al.
Published: (2024)
by: Muqeeth, Mohammed, et al.
Published: (2024)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
by: Patel, Ajay, et al.
Published: (2026)
by: Patel, Ajay, et al.
Published: (2026)
On the Role of Depth in the Expressivity of RNNs
by: Lizaire, Maude, et al.
Published: (2026)
by: Lizaire, Maude, et al.
Published: (2026)
The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
by: Kwok, Devin, et al.
Published: (2025)
by: Kwok, Devin, et al.
Published: (2025)
Free Hunch: Denoiser Covariance Estimation for Diffusion Models Without Extra Costs
by: Rissanen, Severi, et al.
Published: (2024)
by: Rissanen, Severi, et al.
Published: (2024)
UTG: Towards a Unified View of Snapshot and Event Based Models for Temporal Graphs
by: Huang, Shenyang, et al.
Published: (2024)
by: Huang, Shenyang, et al.
Published: (2024)
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
Whoever Started the Interference Should End It: Guiding Data-Free Model Merging via Task Vectors
by: Cheng, Runxi, et al.
Published: (2025)
by: Cheng, Runxi, et al.
Published: (2025)
CPINN-ABPI: Physics-Informed Neural Networks for Accurate Power Estimation in MPSoCs
by: Elshamy, Mohamed R., et al.
Published: (2025)
by: Elshamy, Mohamed R., et al.
Published: (2025)
Private Adaptive Covariance Estimation via Gaussian Graphical Models
by: Ferrando, Cecilia, et al.
Published: (2026)
by: Ferrando, Cecilia, et al.
Published: (2026)
T-GRAB: A Synthetic Diagnostic Benchmark for Learning on Temporal Graphs
by: Dizaji, Alireza, et al.
Published: (2025)
by: Dizaji, Alireza, et al.
Published: (2025)
Generative Learning of Continuous Data by Tensor Networks
by: Meiburg, Alex, et al.
Published: (2023)
by: Meiburg, Alex, et al.
Published: (2023)
Graph Neural Networks for Parameterized Quantum Circuits Expressibility Estimation
by: Aktar, Shamminuj, et al.
Published: (2024)
by: Aktar, Shamminuj, et al.
Published: (2024)
FedMerge: Federated Personalization via Model Merging
by: Chen, Shutong, et al.
Published: (2025)
by: Chen, Shutong, et al.
Published: (2025)
Similar Items
-
Grokking Beyond the Euclidean Norm of Model Parameters
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2025) -
Grokking Finite-Dimensional Algebra
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2026) -
Efficient Probabilistic Tensor Networks
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2025) -
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2024) -
A Tensor Decomposition Perspective on Second-order RNNs
by: Lizaire, Maude, et al.
Published: (2024)