Unified View of Grokking, Double Descent and Emergent Abilities: A Perspective from Circuits Competition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yufei, Hu, Shengding, Han, Xu, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910343074152448
author Huang, Yufei
Hu, Shengding
Han, Xu
Liu, Zhiyuan
Sun, Maosong
author_facet Huang, Yufei
Hu, Shengding
Han, Xu
Liu, Zhiyuan
Sun, Maosong
contents Recent studies have uncovered intriguing phenomena in deep learning, such as grokking, double descent, and emergent abilities in large language models, which challenge human intuition and are crucial for a deeper understanding of neural models. In this paper, we present a comprehensive framework that provides a unified view of these three phenomena, focusing on the competition between memorization and generalization circuits. This approach, initially employed to explain grokking, is extended in our work to encompass a wider range of model sizes and training data volumes. Our framework delineates four distinct training dynamics, each depending on varying combinations of model size and training data quantity. Utilizing this framework, we provide a detailed analysis of the double descent phenomenon and propose two verifiable predictions regarding its occurrence, both substantiated by our experimental results. Moreover, we expand our framework to the multi-task learning paradigm, demonstrating how algorithm tasks can be turned into emergent abilities. This offers a novel perspective to understand emergent abilities in Large Language Models.
format Preprint
id arxiv_https___arxiv_org_abs_2402_15175
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unified View of Grokking, Double Descent and Emergent Abilities: A Perspective from Circuits Competition
Huang, Yufei
Hu, Shengding
Han, Xu
Liu, Zhiyuan
Sun, Maosong
Machine Learning
Recent studies have uncovered intriguing phenomena in deep learning, such as grokking, double descent, and emergent abilities in large language models, which challenge human intuition and are crucial for a deeper understanding of neural models. In this paper, we present a comprehensive framework that provides a unified view of these three phenomena, focusing on the competition between memorization and generalization circuits. This approach, initially employed to explain grokking, is extended in our work to encompass a wider range of model sizes and training data volumes. Our framework delineates four distinct training dynamics, each depending on varying combinations of model size and training data quantity. Utilizing this framework, we provide a detailed analysis of the double descent phenomenon and propose two verifiable predictions regarding its occurrence, both substantiated by our experimental results. Moreover, we expand our framework to the multi-task learning paradigm, demonstrating how algorithm tasks can be turned into emergent abilities. This offers a novel perspective to understand emergent abilities in Large Language Models.
title Unified View of Grokking, Double Descent and Emergent Abilities: A Perspective from Circuits Competition
topic Machine Learning
url https://arxiv.org/abs/2402.15175