DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
Fuente:
arXiv
Saved in:
| Main Authors: | Yoshida, Kotaro, Naraki, Yuji, Horie, Takafumi, Shimizu, Ryotaro, Naganuma, Hiroki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Fairness of Task Arithmetic: The Role of Task Vectors
by: Naganuma, Hiroki, et al.
Published: (2025)
by: Naganuma, Hiroki, et al.
Published: (2025)
Augmenting NER Datasets with LLMs: Towards Automated and Refined Annotation
by: Naraki, Yuji, et al.
Published: (2024)
by: Naraki, Yuji, et al.
Published: (2024)
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
by: Yoshida, Kotaro, et al.
Published: (2024)
by: Yoshida, Kotaro, et al.
Published: (2024)
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
by: Naganuma, Hiroki, et al.
Published: (2023)
by: Naganuma, Hiroki, et al.
Published: (2023)
Revisiting Generalization Measures Beyond IID: An Empirical Study under Distributional Shift
by: Nakai, Sora, et al.
Published: (2026)
by: Nakai, Sora, et al.
Published: (2026)
Robust Invariant Representation Learning by Distribution Extrapolation
by: Yoshida, Kotaro, et al.
Published: (2025)
by: Yoshida, Kotaro, et al.
Published: (2025)
Task Vector Quantization for Memory-Efficient Model Merging
by: Kim, Youngeun, et al.
Published: (2025)
by: Kim, Youngeun, et al.
Published: (2025)
Geometric Insights into Focal Loss: Reducing Curvature for Enhanced Model Calibration
by: Kimura, Masanari, et al.
Published: (2024)
by: Kimura, Masanari, et al.
Published: (2024)
Auto-FlexSwitch: Efficient Dynamic Model Merging via Learnable Task Vector Compression
by: Gao, Junqi, et al.
Published: (2026)
by: Gao, Junqi, et al.
Published: (2026)
Whoever Started the Interference Should End It: Guiding Data-Free Model Merging via Task Vectors
by: Cheng, Runxi, et al.
Published: (2025)
by: Cheng, Runxi, et al.
Published: (2025)
Task Singular Vectors: Reducing Task Interference in Model Merging
by: Gargiulo, Antonio Andrea, et al.
Published: (2024)
by: Gargiulo, Antonio Andrea, et al.
Published: (2024)
Model Merging via Multi-Teacher Knowledge Distillation
by: Dalili, Seyed Arshan, et al.
Published: (2025)
by: Dalili, Seyed Arshan, et al.
Published: (2025)
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
by: Yoshida, Kotaro, et al.
Published: (2025)
by: Yoshida, Kotaro, et al.
Published: (2025)
Convergence Bound and Critical Batch Size of Muon Optimizer
by: Sato, Naoki, et al.
Published: (2025)
by: Sato, Naoki, et al.
Published: (2025)
Scalable Model Merging with Progressive Layer-wise Distillation
by: Xu, Jing, et al.
Published: (2025)
by: Xu, Jing, et al.
Published: (2025)
Robust VAEs via Generating Process of Noise Augmented Data
by: Irobe, Hiroo, et al.
Published: (2024)
by: Irobe, Hiroo, et al.
Published: (2024)
DiLM: Distilling Dataset into Language Model for Text-level Dataset Distillation
by: Maekawa, Aru, et al.
Published: (2024)
by: Maekawa, Aru, et al.
Published: (2024)
Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge Integration
by: Sun, Wenju, et al.
Published: (2025)
by: Sun, Wenju, et al.
Published: (2025)
FedMerge: Federated Personalization via Model Merging
by: Chen, Shutong, et al.
Published: (2025)
by: Chen, Shutong, et al.
Published: (2025)
StatsMerging: Statistics-Guided Model Merging via Task-Specific Teacher Distillation
by: Merugu, Ranjith, et al.
Published: (2025)
by: Merugu, Ranjith, et al.
Published: (2025)
Merge-of-Thought Distillation
by: Shen, Zhanming, et al.
Published: (2025)
by: Shen, Zhanming, et al.
Published: (2025)
No Task Left Behind: Isotropic Model Merging with Common and Task-Specific Subspaces
by: Marczak, Daniel, et al.
Published: (2025)
by: Marczak, Daniel, et al.
Published: (2025)
AdaMerging: Adaptive Model Merging for Multi-Task Learning
by: Yang, Enneng, et al.
Published: (2023)
by: Yang, Enneng, et al.
Published: (2023)
Superpose Task-specific Features for Model Merging
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
Explaining Black-box Model Predictions via Two-level Nested Feature Attributions with Consistency Property
by: Yoshikawa, Yuya, et al.
Published: (2024)
by: Yoshikawa, Yuya, et al.
Published: (2024)
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023)
by: Tam, Derek, et al.
Published: (2023)
Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent
by: Wei, Yongxian, et al.
Published: (2025)
by: Wei, Yongxian, et al.
Published: (2025)
Model Breadcrumbs: Scaling Multi-Task Model Merging with Sparse Masks
by: Davari, MohammadReza, et al.
Published: (2023)
by: Davari, MohammadReza, et al.
Published: (2023)
Less is More: Efficient Model Merging with Binary Task Switch
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
What do near-optimal learning rate schedules look like?
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Toward Theoretical Insights into Diffusion Trajectory Distillation via Operator Merging
by: Gao, Weiguo, et al.
Published: (2025)
by: Gao, Weiguo, et al.
Published: (2025)
Transferring Visual Explainability of Self-Explaining Models to Prediction-Only Models without Additional Training
by: Yoshikawa, Yuya, et al.
Published: (2025)
by: Yoshikawa, Yuya, et al.
Published: (2025)
Merging Multi-Task Models via Weight-Ensembling Mixture of Experts
by: Tang, Anke, et al.
Published: (2024)
by: Tang, Anke, et al.
Published: (2024)
HoRD: Robust Humanoid Control via History-Conditioned Reinforcement Learning and Online Distillation
by: Wang, Puyue, et al.
Published: (2026)
by: Wang, Puyue, et al.
Published: (2026)
Decomposing Task Vectors for Refined Model Editing
by: Damirchi, Hamed, et al.
Published: (2025)
by: Damirchi, Hamed, et al.
Published: (2025)
Mitigating the Backdoor Effect for Multi-Task Model Merging via Safety-Aware Subspace
by: Yang, Jinluan, et al.
Published: (2024)
by: Yang, Jinluan, et al.
Published: (2024)
Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
by: Naganuma, Hiroki, et al.
Published: (2025)
by: Naganuma, Hiroki, et al.
Published: (2025)
Multi-Task Model Merging via Adaptive Weight Disentanglement
by: Xiong, Feng, et al.
Published: (2024)
by: Xiong, Feng, et al.
Published: (2024)
Evaluating Singular Value Thresholds for DNN Weight Matrices based on Random Matrix Theory
by: Nishikawa, Kohei, et al.
Published: (2025)
by: Nishikawa, Kohei, et al.
Published: (2025)
From Task-Specific Models to Unified Systems: A Review of Model Merging Approaches
by: Ruan, Wei, et al.
Published: (2025)
by: Ruan, Wei, et al.
Published: (2025)
Similar Items
-
On Fairness of Task Arithmetic: The Role of Task Vectors
by: Naganuma, Hiroki, et al.
Published: (2025) -
Augmenting NER Datasets with LLMs: Towards Automated and Refined Annotation
by: Naraki, Yuji, et al.
Published: (2024) -
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
by: Yoshida, Kotaro, et al.
Published: (2024) -
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
by: Naganuma, Hiroki, et al.
Published: (2023) -
Revisiting Generalization Measures Beyond IID: An Empirical Study under Distributional Shift
by: Nakai, Sora, et al.
Published: (2026)