Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yan-Lun, Wei, Yi-Ru, Hsu, Chia-Yi, Yu, Chia-Mu, Huang, Chun-Ying, Lin, Ying-Dar, Wu, Yu-Sung, Lee, Wei-Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909515059822592
author Chen, Yan-Lun
Wei, Yi-Ru
Hsu, Chia-Yi
Yu, Chia-Mu
Huang, Chun-Ying
Lin, Ying-Dar
Wu, Yu-Sung
Lee, Wei-Bin
author_facet Chen, Yan-Lun
Wei, Yi-Ru
Hsu, Chia-Yi
Yu, Chia-Mu
Huang, Chun-Ying
Lin, Ying-Dar
Wu, Yu-Sung
Lee, Wei-Bin
contents Large language models (LLMs) demonstrate strong task-specific capabilities through fine-tuning, but merging multiple fine-tuned models often leads to degraded performance due to overlapping instruction-following components. Task Arithmetic (TA), which combines task vectors derived from fine-tuning, enables multi-task learning and task forgetting but struggles to isolate task-specific knowledge from general instruction-following behavior. To address this, we propose Layer-Aware Task Arithmetic (LATA), a novel approach that assigns layer-specific weights to task vectors based on their alignment with instruction-following or task-specific components. By amplifying task-relevant layers and attenuating instruction-following layers, LATA improves task learning and forgetting performance while preserving overall model utility. Experiments on multiple benchmarks, including WikiText-2, GSM8K, and HumanEval, demonstrate that LATA outperforms existing methods in both multi-task learning and selective task forgetting, achieving higher task accuracy and alignment with minimal degradation in output quality. Our findings highlight the importance of layer-wise analysis in disentangling task-specific and general-purpose knowledge, offering a robust framework for efficient model merging and editing.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20186
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
Chen, Yan-Lun
Wei, Yi-Ru
Hsu, Chia-Yi
Yu, Chia-Mu
Huang, Chun-Ying
Lin, Ying-Dar
Wu, Yu-Sung
Lee, Wei-Bin
Computation and Language
Machine Learning
Large language models (LLMs) demonstrate strong task-specific capabilities through fine-tuning, but merging multiple fine-tuned models often leads to degraded performance due to overlapping instruction-following components. Task Arithmetic (TA), which combines task vectors derived from fine-tuning, enables multi-task learning and task forgetting but struggles to isolate task-specific knowledge from general instruction-following behavior. To address this, we propose Layer-Aware Task Arithmetic (LATA), a novel approach that assigns layer-specific weights to task vectors based on their alignment with instruction-following or task-specific components. By amplifying task-relevant layers and attenuating instruction-following layers, LATA improves task learning and forgetting performance while preserving overall model utility. Experiments on multiple benchmarks, including WikiText-2, GSM8K, and HumanEval, demonstrate that LATA outperforms existing methods in both multi-task learning and selective task forgetting, achieving higher task accuracy and alignment with minimal degradation in output quality. Our findings highlight the importance of layer-wise analysis in disentangling task-specific and general-purpose knowledge, offering a robust framework for efficient model merging and editing.
title Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.20186