Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917802215997440 |
|---|---|
| author | Aakanksha Ahmadian, Arash Goldfarb-Tarrant, Seraphina Ermis, Beyza Fadaee, Marzieh Hooker, Sara |
| author_facet | Aakanksha Ahmadian, Arash Goldfarb-Tarrant, Seraphina Ermis, Beyza Fadaee, Marzieh Hooker, Sara |
| contents | Large Language Models (LLMs) have been adopted and deployed worldwide for a broad variety of applications. However, ensuring their safe use remains a significant challenge. Preference training and safety measures often overfit to harms prevalent in Western-centric datasets, and safety protocols frequently fail to extend to multilingual settings. In this work, we explore model merging in a diverse multi-task setting, combining safety and general-purpose tasks within a multilingual context. Each language introduces unique and varied learning challenges across tasks. We find that objective-based merging is more effective than mixing data, with improvements of up to 8% and 10% in general performance and safety respectively. We also find that language-based merging is highly effective -- by merging monolingually fine-tuned models, we achieve a 4% increase in general performance and 7% reduction in harm across all languages on top of the data mixtures method using the same available data. Overall, our comprehensive study of merging approaches provides a useful framework for building strong and safe multilingual models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_10801 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning Aakanksha Ahmadian, Arash Goldfarb-Tarrant, Seraphina Ermis, Beyza Fadaee, Marzieh Hooker, Sara Computation and Language Machine Learning Large Language Models (LLMs) have been adopted and deployed worldwide for a broad variety of applications. However, ensuring their safe use remains a significant challenge. Preference training and safety measures often overfit to harms prevalent in Western-centric datasets, and safety protocols frequently fail to extend to multilingual settings. In this work, we explore model merging in a diverse multi-task setting, combining safety and general-purpose tasks within a multilingual context. Each language introduces unique and varied learning challenges across tasks. We find that objective-based merging is more effective than mixing data, with improvements of up to 8% and 10% in general performance and safety respectively. We also find that language-based merging is highly effective -- by merging monolingually fine-tuned models, we achieve a 4% increase in general performance and 7% reduction in harm across all languages on top of the data mixtures method using the same available data. Overall, our comprehensive study of merging approaches provides a useful framework for building strong and safe multilingual models. |
| title | Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2410.10801 |