Model Merging via Data-Free Covariance Estimation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hameed, Marawan Gamal Abdel, Tam, Derek, Notsawo, Pascal Jr Tikeng, Raffel, Colin, Rabusseau, Guillaume |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Grokking Beyond the Euclidean Norm of Model Parameters
par: Notsawo, Pascal Jr Tikeng, et autres
Publié: (2025)
par: Notsawo, Pascal Jr Tikeng, et autres
Publié: (2025)
Grokking Finite-Dimensional Algebra
par: Notsawo, Pascal Jr Tikeng, et autres
Publié: (2026)
par: Notsawo, Pascal Jr Tikeng, et autres
Publié: (2026)
Efficient Probabilistic Tensor Networks
par: Hameed, Marawan Gamal Abdel, et autres
Publié: (2025)
par: Hameed, Marawan Gamal Abdel, et autres
Publié: (2025)
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
par: Hameed, Marawan Gamal Abdel, et autres
Publié: (2024)
par: Hameed, Marawan Gamal Abdel, et autres
Publié: (2024)
A Tensor Decomposition Perspective on Second-order RNNs
par: Lizaire, Maude, et autres
Publié: (2024)
par: Lizaire, Maude, et autres
Publié: (2024)
Merging by Matching Models in Task Parameter Subspaces
par: Tam, Derek, et autres
Publié: (2023)
par: Tam, Derek, et autres
Publié: (2023)
Realistic Evaluation of Model Merging for Compositional Generalization
par: Tam, Derek, et autres
Publié: (2024)
par: Tam, Derek, et autres
Publié: (2024)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
par: Li, YuXin, et autres
Publié: (2025)
par: Li, YuXin, et autres
Publié: (2025)
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
par: Je, Gyung Hyun, et autres
Publié: (2025)
par: Je, Gyung Hyun, et autres
Publié: (2025)
Soft Merging of Experts with Adaptive Routing
par: Muqeeth, Mohammed, et autres
Publié: (2023)
par: Muqeeth, Mohammed, et autres
Publié: (2023)
Lost in Translation: The Algorithmic Gap Between LMs and the Brain
par: Tosato, Tommaso, et autres
Publié: (2024)
par: Tosato, Tommaso, et autres
Publié: (2024)
Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
par: Matena, Michael, et autres
Publié: (2023)
par: Matena, Michael, et autres
Publié: (2023)
Tractable Shapley Values and Interactions via Tensor Networks
par: Heidari, Farzaneh, et autres
Publié: (2025)
par: Heidari, Farzaneh, et autres
Publié: (2025)
Tensor Cookbook: Mastering Tensors through Diagrams
par: Rakhshan, Beheshteh T., et autres
Publié: (2026)
par: Rakhshan, Beheshteh T., et autres
Publié: (2026)
TN-SHAP-G: Graph-Structured Tensor Network Surrogates for Shapley Values and Interactions
par: Heidari, Farzaneh, et autres
Publié: (2026)
par: Heidari, Farzaneh, et autres
Publié: (2026)
Position: The Most Expensive Part of an LLM should be its Training Data
par: Kandpal, Nikhil, et autres
Publié: (2025)
par: Kandpal, Nikhil, et autres
Publié: (2025)
Higher Order Transformers: Enhancing Stock Movement Prediction On Multimodal Time-Series Data
par: Omranpour, Soroush, et autres
Publié: (2024)
par: Omranpour, Soroush, et autres
Publié: (2024)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
par: Patel, Ajay, et autres
Publié: (2024)
par: Patel, Ajay, et autres
Publié: (2024)
KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
par: Lesens, Damien, et autres
Publié: (2025)
par: Lesens, Damien, et autres
Publié: (2025)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
par: Rizvi-Martel, Michael, et autres
Publié: (2026)
par: Rizvi-Martel, Michael, et autres
Publié: (2026)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
par: Liu, Fengyuan, et autres
Publié: (2024)
par: Liu, Fengyuan, et autres
Publié: (2024)
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
par: Liu, Haokun, et autres
Publié: (2026)
par: Liu, Haokun, et autres
Publié: (2026)
Higher-Order Transformers With Kronecker-Structured Attention
par: Omranpour, Soroush, et autres
Publié: (2024)
par: Omranpour, Soroush, et autres
Publié: (2024)
Enhancing Training Data Attribution with Representational Optimization
par: Sun, Weiwei, et autres
Publié: (2025)
par: Sun, Weiwei, et autres
Publié: (2025)
FlowQ-Net: A Generative Framework for Automated Quantum Circuit Design
par: Dai, Jun, et autres
Publié: (2025)
par: Dai, Jun, et autres
Publié: (2025)
Learning to Route Among Specialized Experts for Zero-Shot Generalization
par: Muqeeth, Mohammed, et autres
Publié: (2024)
par: Muqeeth, Mohammed, et autres
Publié: (2024)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
par: Yadav, Prateek, et autres
Publié: (2023)
par: Yadav, Prateek, et autres
Publié: (2023)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
par: Patel, Ajay, et autres
Publié: (2026)
par: Patel, Ajay, et autres
Publié: (2026)
On the Role of Depth in the Expressivity of RNNs
par: Lizaire, Maude, et autres
Publié: (2026)
par: Lizaire, Maude, et autres
Publié: (2026)
The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
par: Kwok, Devin, et autres
Publié: (2025)
par: Kwok, Devin, et autres
Publié: (2025)
Free Hunch: Denoiser Covariance Estimation for Diffusion Models Without Extra Costs
par: Rissanen, Severi, et autres
Publié: (2024)
par: Rissanen, Severi, et autres
Publié: (2024)
UTG: Towards a Unified View of Snapshot and Event Based Models for Temporal Graphs
par: Huang, Shenyang, et autres
Publié: (2024)
par: Huang, Shenyang, et autres
Publié: (2024)
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
par: Si, Chongjie, et autres
Publié: (2025)
par: Si, Chongjie, et autres
Publié: (2025)
Whoever Started the Interference Should End It: Guiding Data-Free Model Merging via Task Vectors
par: Cheng, Runxi, et autres
Publié: (2025)
par: Cheng, Runxi, et autres
Publié: (2025)
CPINN-ABPI: Physics-Informed Neural Networks for Accurate Power Estimation in MPSoCs
par: Elshamy, Mohamed R., et autres
Publié: (2025)
par: Elshamy, Mohamed R., et autres
Publié: (2025)
Private Adaptive Covariance Estimation via Gaussian Graphical Models
par: Ferrando, Cecilia, et autres
Publié: (2026)
par: Ferrando, Cecilia, et autres
Publié: (2026)
T-GRAB: A Synthetic Diagnostic Benchmark for Learning on Temporal Graphs
par: Dizaji, Alireza, et autres
Publié: (2025)
par: Dizaji, Alireza, et autres
Publié: (2025)
Generative Learning of Continuous Data by Tensor Networks
par: Meiburg, Alex, et autres
Publié: (2023)
par: Meiburg, Alex, et autres
Publié: (2023)
Graph Neural Networks for Parameterized Quantum Circuits Expressibility Estimation
par: Aktar, Shamminuj, et autres
Publié: (2024)
par: Aktar, Shamminuj, et autres
Publié: (2024)
FedMerge: Federated Personalization via Model Merging
par: Chen, Shutong, et autres
Publié: (2025)
par: Chen, Shutong, et autres
Publié: (2025)
Documents similaires
-
Grokking Beyond the Euclidean Norm of Model Parameters
par: Notsawo, Pascal Jr Tikeng, et autres
Publié: (2025) -
Grokking Finite-Dimensional Algebra
par: Notsawo, Pascal Jr Tikeng, et autres
Publié: (2026) -
Efficient Probabilistic Tensor Networks
par: Hameed, Marawan Gamal Abdel, et autres
Publié: (2025) -
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
par: Hameed, Marawan Gamal Abdel, et autres
Publié: (2024) -
A Tensor Decomposition Perspective on Second-order RNNs
par: Lizaire, Maude, et autres
Publié: (2024)