Scaling Laws for Linear Complexity Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Xuyang, Li, Dong, Leng, Ruitao, Qin, Zhen, Sun, Weigao, Zhong, Yiran
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909230314815488
author Shen, Xuyang
Li, Dong
Leng, Ruitao
Qin, Zhen
Sun, Weigao
Zhong, Yiran
author_facet Shen, Xuyang
Li, Dong
Leng, Ruitao
Qin, Zhen
Sun, Weigao
Zhong, Yiran
contents The interest in linear complexity models for large language models is on the rise, although their scaling capacity remains uncertain. In this study, we present the scaling laws for linear complexity language models to establish a foundation for their scalability. Specifically, we examine the scaling behaviors of three efficient linear architectures. These include TNL, a linear attention model with data-independent decay; HGRN2, a linear RNN with data-dependent decay; and cosFormer2, a linear attention model without decay. We also include LLaMA as a baseline architecture for softmax attention for comparison. These models were trained with six variants, ranging from 70M to 7B parameters on a 300B-token corpus, and evaluated with a total of 1,376 intermediate checkpoints on various downstream tasks. These tasks include validation loss, commonsense reasoning, and information retrieval and generation. The study reveals that existing linear complexity language models exhibit similar scaling capabilities as conventional transformer-based models while also demonstrating superior linguistic proficiency and knowledge retention.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16690
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Laws for Linear Complexity Language Models
Shen, Xuyang
Li, Dong
Leng, Ruitao
Qin, Zhen
Sun, Weigao
Zhong, Yiran
Computation and Language
The interest in linear complexity models for large language models is on the rise, although their scaling capacity remains uncertain. In this study, we present the scaling laws for linear complexity language models to establish a foundation for their scalability. Specifically, we examine the scaling behaviors of three efficient linear architectures. These include TNL, a linear attention model with data-independent decay; HGRN2, a linear RNN with data-dependent decay; and cosFormer2, a linear attention model without decay. We also include LLaMA as a baseline architecture for softmax attention for comparison. These models were trained with six variants, ranging from 70M to 7B parameters on a 300B-token corpus, and evaluated with a total of 1,376 intermediate checkpoints on various downstream tasks. These tasks include validation loss, commonsense reasoning, and information retrieval and generation. The study reveals that existing linear complexity language models exhibit similar scaling capabilities as conventional transformer-based models while also demonstrating superior linguistic proficiency and knowledge retention.
title Scaling Laws for Linear Complexity Language Models
topic Computation and Language
url https://arxiv.org/abs/2406.16690