Harnessing Optimization Dynamics for Curvature-Informed Model Merging
Fuente:
arXiv
Saved in:
| Main Authors: | Mahdavinia, Pouria, Mahdavi, Hamed, Mireshghallah, Niloofar, Mahdavi, Mehrdad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Low-rank Momentum Factorization for Memory Efficient Training
by: Mahdavinia, Pouria, et al.
Published: (2025)
by: Mahdavinia, Pouria, et al.
Published: (2025)
RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
by: Mahdavi, Hamed, et al.
Published: (2025)
by: Mahdavi, Hamed, et al.
Published: (2025)
Model Merging via Multi-Teacher Knowledge Distillation
by: Dalili, Seyed Arshan, et al.
Published: (2025)
by: Dalili, Seyed Arshan, et al.
Published: (2025)
Merge before Forget: A Single LoRA Continual Learning via Continual Merging
by: Qiao, Fuli, et al.
Published: (2025)
by: Qiao, Fuli, et al.
Published: (2025)
On the Generalization Capability of Temporal Graph Learning Algorithms: Theoretical Insights and a Simpler Method
by: Cong, Weilin, et al.
Published: (2024)
by: Cong, Weilin, et al.
Published: (2024)
Operationalizing Data Minimization for Privacy-Preserving LLM Prompting
by: Zhou, Jijie, et al.
Published: (2025)
by: Zhou, Jijie, et al.
Published: (2025)
Position: Privacy Is Not Just Memorization!
by: Mireshghallah, Niloofar, et al.
Published: (2025)
by: Mireshghallah, Niloofar, et al.
Published: (2025)
SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?
by: Han, Kevin, et al.
Published: (2026)
by: Han, Kevin, et al.
Published: (2026)
On Large-scale Evaluation of Embedding Models for Knowledge Graph Completion
by: Shirvani-Mahdavi, Nasim, et al.
Published: (2025)
by: Shirvani-Mahdavi, Nasim, et al.
Published: (2025)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
by: Thrampoulidis, Christos, et al.
Published: (2025)
by: Thrampoulidis, Christos, et al.
Published: (2025)
Lotus at SemEval-2025 Task 11: RoBERTa with Llama-3 Generated Explanations for Multi-Label Emotion Classification
by: Ranjbar, Niloofar, et al.
Published: (2025)
by: Ranjbar, Niloofar, et al.
Published: (2025)
Differentially Private Learning Needs Better Model Initialization and Self-Distillation
by: Ngong, Ivoline C., et al.
Published: (2024)
by: Ngong, Ivoline C., et al.
Published: (2024)
Causal Unlearning in Collaborative Optimization: Exact and Approximate Influence Reversal under Adversarial Contributions
by: Mahdavi, Ali, et al.
Published: (2026)
by: Mahdavi, Ali, et al.
Published: (2026)
CombiGraph-Vis: A Curated Multimodal Olympiad Benchmark for Discrete Mathematical Reasoning
by: Mahdavi, Hamed, et al.
Published: (2025)
by: Mahdavi, Hamed, et al.
Published: (2025)
From Graph Diffusion to Graph Classification
by: Xian, Jia Jun Cheng, et al.
Published: (2024)
by: Xian, Jia Jun Cheng, et al.
Published: (2024)
Learning Stable Predictors from Weak Supervision under Distribution Shift
by: Shoeibi, Mehrdad, et al.
Published: (2026)
by: Shoeibi, Mehrdad, et al.
Published: (2026)
Stochastic Compositional Minimax Optimization with Provable Convergence Guarantees
by: Deng, Yuyang, et al.
Published: (2024)
by: Deng, Yuyang, et al.
Published: (2024)
Brains vs. Bytes: Evaluating LLM Proficiency in Olympiad Mathematics
by: Mahdavi, Hamed, et al.
Published: (2025)
by: Mahdavi, Hamed, et al.
Published: (2025)
Feature-Function Curvature Analysis: A Geometric Framework for Explaining Differentiable Models
by: Najafi, Hamed, et al.
Published: (2025)
by: Najafi, Hamed, et al.
Published: (2025)
Rule2Text: Natural Language Explanation of Logical Rules in Knowledge Graphs
by: Shirvani-Mahdavi, Nasim, et al.
Published: (2025)
by: Shirvani-Mahdavi, Nasim, et al.
Published: (2025)
FW-Merging: Scaling Model Merging with Frank-Wolfe Optimization
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
PSO-Merging: Merging Models Based on Particle Swarm Optimization
by: Zhang, Kehao, et al.
Published: (2025)
by: Zhang, Kehao, et al.
Published: (2025)
Integrating Large Language Models in Financial Investments and Market Analysis: A Survey
by: Mahdavi, Sedigheh, et al.
Published: (2025)
by: Mahdavi, Sedigheh, et al.
Published: (2025)
BIOGEN: Evidence-Grounded Multi-Agent Reasoning Framework for Transcriptomic Interpretation in Antimicrobial Resistance
by: Hossain, Elias, et al.
Published: (2025)
by: Hossain, Elias, et al.
Published: (2025)
On the Convergence and Stability of Distributed Sub-model Training
by: Deng, Yuyang, et al.
Published: (2025)
by: Deng, Yuyang, et al.
Published: (2025)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
by: Chen, Tong, et al.
Published: (2025)
by: Chen, Tong, et al.
Published: (2025)
MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model Merging
by: Wang, Jiapeng, et al.
Published: (2026)
by: Wang, Jiapeng, et al.
Published: (2026)
BD-Merging: Bias-Aware Dynamic Model Merging with Evidence-Guided Contrastive Learning
by: Xie, Yuhan, et al.
Published: (2026)
by: Xie, Yuhan, et al.
Published: (2026)
MIN-Merging: Merge the Important Neurons for Model Merging
by: Liang, Yunfei
Published: (2025)
by: Liang, Yunfei
Published: (2025)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
by: Mahdavi, Sadegh, et al.
Published: (2025)
by: Mahdavi, Sadegh, et al.
Published: (2025)
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
by: Lu, Zhenyi, et al.
Published: (2024)
by: Lu, Zhenyi, et al.
Published: (2024)
MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
Self-Supervised Image Super-Resolution Quality Assessment based on Content-Free Multi-Model Oriented Representation Learning
by: Majlessi, Kian, et al.
Published: (2026)
by: Majlessi, Kian, et al.
Published: (2026)
HARBOR: Automated Harness Optimization
by: Sengupta, Biswa, et al.
Published: (2026)
by: Sengupta, Biswa, et al.
Published: (2026)
Dynamic Model Merging Made Slim
by: Du, Guodong, et al.
Published: (2026)
by: Du, Guodong, et al.
Published: (2026)
A New Approach to Backtracking Counterfactual Explanations: A Unified Causal Framework for Efficient Model Interpretability
by: Fatemi, Pouria, et al.
Published: (2025)
by: Fatemi, Pouria, et al.
Published: (2025)
Bayesian Model Merging
by: Li, Kaiyang, et al.
Published: (2026)
by: Li, Kaiyang, et al.
Published: (2026)
Quantum Speedups for Markov Chain Monte Carlo Methods with Application to Optimization
by: Ozgul, Guneykan, et al.
Published: (2025)
by: Ozgul, Guneykan, et al.
Published: (2025)
Merging Smarter, Generalizing Better: Enhancing Model Merging on OOD Data
by: Zhang, Bingjie, et al.
Published: (2025)
by: Zhang, Bingjie, et al.
Published: (2025)
Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking
by: Chaichana, Yuatyong, et al.
Published: (2025)
by: Chaichana, Yuatyong, et al.
Published: (2025)
Similar Items
-
Low-rank Momentum Factorization for Memory Efficient Training
by: Mahdavinia, Pouria, et al.
Published: (2025) -
RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
by: Mahdavi, Hamed, et al.
Published: (2025) -
Model Merging via Multi-Teacher Knowledge Distillation
by: Dalili, Seyed Arshan, et al.
Published: (2025) -
Merge before Forget: A Single LoRA Continual Learning via Continual Merging
by: Qiao, Fuli, et al.
Published: (2025) -
On the Generalization Capability of Temporal Graph Learning Algorithms: Theoretical Insights and a Simpler Method
by: Cong, Weilin, et al.
Published: (2024)