Hierarchical vs. Flat Iteration in Shared-Weight Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Han, Sang-Il
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915939492036608
author Han, Sang-Il
author_facet Han, Sang-Il
contents We present an empirical study of whether hierarchically structured, shared-weight recurrence can match the representational quality of independent-layer stacking in a Transformer-based language model. HRM-LM replaces L independent Transformer layers with a two-speed recurrent pair: a Fast module operating at every step for local refinement, and a Slow module operating every T steps for global compression. This recurrent hierarchy is unrolled for M = N x T steps with shared parameters. The central and most robust finding, supported by a parameter-matched Universal Transformer ablation (UniTF, 1.2B) across five independent runs, is a sharp empirical gap between the two approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14442
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hierarchical vs. Flat Iteration in Shared-Weight Transformers
Han, Sang-Il
Computation and Language
Artificial Intelligence
We present an empirical study of whether hierarchically structured, shared-weight recurrence can match the representational quality of independent-layer stacking in a Transformer-based language model. HRM-LM replaces L independent Transformer layers with a two-speed recurrent pair: a Fast module operating at every step for local refinement, and a Slow module operating every T steps for global compression. This recurrent hierarchy is unrolled for M = N x T steps with shared parameters. The central and most robust finding, supported by a parameter-matched Universal Transformer ablation (UniTF, 1.2B) across five independent runs, is a sharp empirical gap between the two approaches.
title Hierarchical vs. Flat Iteration in Shared-Weight Transformers
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.14442