Predicting Emergent Abilities with Infinite Resolution Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Shengding, Liu, Xin, Han, Xu, Zhang, Xinrong, He, Chaoqun, Zhao, Weilin, Lin, Yankai, Ding, Ning, Ou, Zebin, Zeng, Guoyang, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911842640592896
author Hu, Shengding
Liu, Xin
Han, Xu
Zhang, Xinrong
He, Chaoqun
Zhao, Weilin
Lin, Yankai
Ding, Ning
Ou, Zebin
Zeng, Guoyang
Liu, Zhiyuan
Sun, Maosong
author_facet Hu, Shengding
Liu, Xin
Han, Xu
Zhang, Xinrong
He, Chaoqun
Zhao, Weilin
Lin, Yankai
Ding, Ning
Ou, Zebin
Zeng, Guoyang
Liu, Zhiyuan
Sun, Maosong
contents The scientific scale-up of large language models (LLMs) necessitates a comprehensive understanding of their scaling properties. However, the existing literature on the scaling properties only yields an incomplete answer: optimization loss decreases predictably as the model size increases, in line with established scaling law; yet no scaling law for task has been established and the task performances are far from predictable during scaling. Task performances typically show minor gains on small models until they improve dramatically once models exceed a size threshold, exemplifying the ``emergent abilities''. In this study, we discover that small models, although they exhibit minor performance, demonstrate critical and consistent task performance improvements that are not captured by conventional evaluation strategies due to insufficient measurement resolution. To measure such improvements, we introduce PassUntil, an evaluation strategy with theoretically infinite resolution, through massive sampling in the decoding phase. With PassUntil, we conduct a quantitative investigation into the scaling law of task performance. The investigation contains two parts. Firstly, a strict task scaling law that is not conventionally known to exist, is identified, enhancing the predictability of task performances. Remarkably, we are able to predict the performance of the 2.4B model on code generation with merely 0.05\% deviation before training starts, which is the first systematic attempt to verify predictable scaling proposed by GPT-4's report. Secondly, we are able to study emergent abilities quantitatively. We identify a kind of accelerated emergence whose scaling curve cannot be fitted by standard scaling law function and has a increasing speed. We then examine two hypothesis and imply that the ``multiple circuits hypothesis'' might be responsible for the accelerated emergence.
format Preprint
id arxiv_https___arxiv_org_abs_2310_03262
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Predicting Emergent Abilities with Infinite Resolution Evaluation
Hu, Shengding
Liu, Xin
Han, Xu
Zhang, Xinrong
He, Chaoqun
Zhao, Weilin
Lin, Yankai
Ding, Ning
Ou, Zebin
Zeng, Guoyang
Liu, Zhiyuan
Sun, Maosong
Computation and Language
The scientific scale-up of large language models (LLMs) necessitates a comprehensive understanding of their scaling properties. However, the existing literature on the scaling properties only yields an incomplete answer: optimization loss decreases predictably as the model size increases, in line with established scaling law; yet no scaling law for task has been established and the task performances are far from predictable during scaling. Task performances typically show minor gains on small models until they improve dramatically once models exceed a size threshold, exemplifying the ``emergent abilities''. In this study, we discover that small models, although they exhibit minor performance, demonstrate critical and consistent task performance improvements that are not captured by conventional evaluation strategies due to insufficient measurement resolution. To measure such improvements, we introduce PassUntil, an evaluation strategy with theoretically infinite resolution, through massive sampling in the decoding phase. With PassUntil, we conduct a quantitative investigation into the scaling law of task performance. The investigation contains two parts. Firstly, a strict task scaling law that is not conventionally known to exist, is identified, enhancing the predictability of task performances. Remarkably, we are able to predict the performance of the 2.4B model on code generation with merely 0.05\% deviation before training starts, which is the first systematic attempt to verify predictable scaling proposed by GPT-4's report. Secondly, we are able to study emergent abilities quantitatively. We identify a kind of accelerated emergence whose scaling curve cannot be fitted by standard scaling law function and has a increasing speed. We then examine two hypothesis and imply that the ``multiple circuits hypothesis'' might be responsible for the accelerated emergence.
title Predicting Emergent Abilities with Infinite Resolution Evaluation
topic Computation and Language
url https://arxiv.org/abs/2310.03262