Deeper Insights Without Updates: The Power of In-Context Learning Over Fine-Tuning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yin, Qingyu, He, Xuzheng, Deng, Luoao, Leong, Chak Tou, Wang, Fan, Yan, Yanzhao, Shen, Xiaoyu, Zhang, Qiang
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913535434424320
author Yin, Qingyu
He, Xuzheng
Deng, Luoao
Leong, Chak Tou
Wang, Fan
Yan, Yanzhao
Shen, Xiaoyu
Zhang, Qiang
author_facet Yin, Qingyu
He, Xuzheng
Deng, Luoao
Leong, Chak Tou
Wang, Fan
Yan, Yanzhao
Shen, Xiaoyu
Zhang, Qiang
contents Fine-tuning and in-context learning (ICL) are two prevalent methods in imbuing large language models with task-specific knowledge. It is commonly believed that fine-tuning can surpass ICL given sufficient training samples as it allows the model to adjust its internal parameters based on the data. However, this paper presents a counterintuitive finding: For tasks with implicit patterns, ICL captures these patterns significantly better than fine-tuning. We developed several datasets featuring implicit patterns, such as sequences determining answers through parity or identifying reducible terms in calculations. We then evaluated the models' understanding of these patterns under both fine-tuning and ICL across models ranging from 0.5B to 7B parameters. The results indicate that models employing ICL can quickly grasp deep patterns and significantly improve accuracy. In contrast, fine-tuning, despite utilizing thousands of times more training samples than ICL, achieved only limited improvements. We also proposed circuit shift theory from a mechanistic interpretability's view to explain why ICL wins.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04691
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Deeper Insights Without Updates: The Power of In-Context Learning Over Fine-Tuning
Yin, Qingyu
He, Xuzheng
Deng, Luoao
Leong, Chak Tou
Wang, Fan
Yan, Yanzhao
Shen, Xiaoyu
Zhang, Qiang
Machine Learning
Computation and Language
Fine-tuning and in-context learning (ICL) are two prevalent methods in imbuing large language models with task-specific knowledge. It is commonly believed that fine-tuning can surpass ICL given sufficient training samples as it allows the model to adjust its internal parameters based on the data. However, this paper presents a counterintuitive finding: For tasks with implicit patterns, ICL captures these patterns significantly better than fine-tuning. We developed several datasets featuring implicit patterns, such as sequences determining answers through parity or identifying reducible terms in calculations. We then evaluated the models' understanding of these patterns under both fine-tuning and ICL across models ranging from 0.5B to 7B parameters. The results indicate that models employing ICL can quickly grasp deep patterns and significantly improve accuracy. In contrast, fine-tuning, despite utilizing thousands of times more training samples than ICL, achieved only limited improvements. We also proposed circuit shift theory from a mechanistic interpretability's view to explain why ICL wins.
title Deeper Insights Without Updates: The Power of In-Context Learning Over Fine-Tuning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2410.04691