Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haller, Patrick, Golde, Jonas, Akbik, Alan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908640011616256
author Haller, Patrick
Golde, Jonas
Akbik, Alan
author_facet Haller, Patrick
Golde, Jonas
Akbik, Alan
contents We study architectural and optimization techniques for sample-efficient language modeling under the constraints of the BabyLM 2025 shared task. Our model, BLaLM, replaces self-attention with a linear-time mLSTM token mixer and explores lightweight enhancements, including short convolutions, sliding window attention with dynamic modulation, and Hedgehog feature maps. To support training in low-resource settings, we curate a high-quality corpus emphasizing readability and pedagogical structure. Experiments across both STRICT and STRICT-SMALL tracks show that (1) linear attention combined with sliding window attention consistently improves zero-shot performance, and (2) the Muon optimizer stabilizes convergence and reduces perplexity over AdamW. These results highlight effective strategies for efficient language modeling without relying on scale.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05560
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
Haller, Patrick
Golde, Jonas
Akbik, Alan
Computation and Language
Artificial Intelligence
We study architectural and optimization techniques for sample-efficient language modeling under the constraints of the BabyLM 2025 shared task. Our model, BLaLM, replaces self-attention with a linear-time mLSTM token mixer and explores lightweight enhancements, including short convolutions, sliding window attention with dynamic modulation, and Hedgehog feature maps. To support training in low-resource settings, we curate a high-quality corpus emphasizing readability and pedagogical structure. Experiments across both STRICT and STRICT-SMALL tracks show that (1) linear attention combined with sliding window attention consistently improves zero-shot performance, and (2) the Muon optimizer stabilizes convergence and reduces perplexity over AdamW. These results highlight effective strategies for efficient language modeling without relying on scale.
title Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.05560