Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909515059822592 |
|---|---|
| author | Chen, Yan-Lun Wei, Yi-Ru Hsu, Chia-Yi Yu, Chia-Mu Huang, Chun-Ying Lin, Ying-Dar Wu, Yu-Sung Lee, Wei-Bin |
| author_facet | Chen, Yan-Lun Wei, Yi-Ru Hsu, Chia-Yi Yu, Chia-Mu Huang, Chun-Ying Lin, Ying-Dar Wu, Yu-Sung Lee, Wei-Bin |
| contents | Large language models (LLMs) demonstrate strong task-specific capabilities through fine-tuning, but merging multiple fine-tuned models often leads to degraded performance due to overlapping instruction-following components. Task Arithmetic (TA), which combines task vectors derived from fine-tuning, enables multi-task learning and task forgetting but struggles to isolate task-specific knowledge from general instruction-following behavior. To address this, we propose Layer-Aware Task Arithmetic (LATA), a novel approach that assigns layer-specific weights to task vectors based on their alignment with instruction-following or task-specific components. By amplifying task-relevant layers and attenuating instruction-following layers, LATA improves task learning and forgetting performance while preserving overall model utility. Experiments on multiple benchmarks, including WikiText-2, GSM8K, and HumanEval, demonstrate that LATA outperforms existing methods in both multi-task learning and selective task forgetting, achieving higher task accuracy and alignment with minimal degradation in output quality. Our findings highlight the importance of layer-wise analysis in disentangling task-specific and general-purpose knowledge, offering a robust framework for efficient model merging and editing. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_20186 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge Chen, Yan-Lun Wei, Yi-Ru Hsu, Chia-Yi Yu, Chia-Mu Huang, Chun-Ying Lin, Ying-Dar Wu, Yu-Sung Lee, Wei-Bin Computation and Language Machine Learning Large language models (LLMs) demonstrate strong task-specific capabilities through fine-tuning, but merging multiple fine-tuned models often leads to degraded performance due to overlapping instruction-following components. Task Arithmetic (TA), which combines task vectors derived from fine-tuning, enables multi-task learning and task forgetting but struggles to isolate task-specific knowledge from general instruction-following behavior. To address this, we propose Layer-Aware Task Arithmetic (LATA), a novel approach that assigns layer-specific weights to task vectors based on their alignment with instruction-following or task-specific components. By amplifying task-relevant layers and attenuating instruction-following layers, LATA improves task learning and forgetting performance while preserving overall model utility. Experiments on multiple benchmarks, including WikiText-2, GSM8K, and HumanEval, demonstrate that LATA outperforms existing methods in both multi-task learning and selective task forgetting, achieving higher task accuracy and alignment with minimal degradation in output quality. Our findings highlight the importance of layer-wise analysis in disentangling task-specific and general-purpose knowledge, offering a robust framework for efficient model merging and editing. |
| title | Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2502.20186 |