Understanding Emergent Abilities of Language Models from the Loss Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Zhengxiao, Zeng, Aohan, Dong, Yuxiao, Tang, Jie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913650260836352
author Du, Zhengxiao
Zeng, Aohan
Dong, Yuxiao
Tang, Jie
author_facet Du, Zhengxiao
Zeng, Aohan
Dong, Yuxiao
Tang, Jie
contents Recent studies have put into question the belief that emergent abilities in language models are exclusive to large models. This skepticism arises from two observations: 1) smaller models can also exhibit high performance on emergent abilities and 2) there is doubt on the discontinuous metrics used to measure these abilities. In this paper, we propose to study emergent abilities in the lens of pre-training loss, instead of model size or training compute. We demonstrate that the Transformer models with the same pre-training loss, but different model and data sizes, generate the same performance on various downstream tasks, with a fixed data corpus, tokenization, and model architecture. We also discover that a model exhibits emergent abilities on certain tasks -- regardless of the continuity of metrics -- when its pre-training loss falls below a specific threshold. Before reaching this threshold, its performance remains at the level of random guessing. This inspires us to redefine emergent abilities as those that manifest in models with lower pre-training losses, highlighting that these abilities cannot be predicted by merely extrapolating the performance trends of models with higher pre-training losses.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15796
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding Emergent Abilities of Language Models from the Loss Perspective
Du, Zhengxiao
Zeng, Aohan
Dong, Yuxiao
Tang, Jie
Computation and Language
Artificial Intelligence
Machine Learning
Recent studies have put into question the belief that emergent abilities in language models are exclusive to large models. This skepticism arises from two observations: 1) smaller models can also exhibit high performance on emergent abilities and 2) there is doubt on the discontinuous metrics used to measure these abilities. In this paper, we propose to study emergent abilities in the lens of pre-training loss, instead of model size or training compute. We demonstrate that the Transformer models with the same pre-training loss, but different model and data sizes, generate the same performance on various downstream tasks, with a fixed data corpus, tokenization, and model architecture. We also discover that a model exhibits emergent abilities on certain tasks -- regardless of the continuity of metrics -- when its pre-training loss falls below a specific threshold. Before reaching this threshold, its performance remains at the level of random guessing. This inspires us to redefine emergent abilities as those that manifest in models with lower pre-training losses, highlighting that these abilities cannot be predicted by merely extrapolating the performance trends of models with higher pre-training losses.
title Understanding Emergent Abilities of Language Models from the Loss Perspective
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2403.15796