What Makes Looped Transformers Perform Better Than Non-Recursive Ones

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gong, Zixuan, Liu, Yong, Teng, Jiaye
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917185468760064
author Gong, Zixuan
Liu, Yong
Teng, Jiaye
author_facet Gong, Zixuan
Liu, Yong
Teng, Jiaye
contents While looped transformers (termed as Looped-Attn) often outperform standard transformers (termed as Single-Attn) on complex reasoning tasks, the mechanism for this advantage remains underexplored. In this paper, we explain this phenomenon through the lens of loss landscape geometry, inspired by empirical observations of their distinct dynamics at both sample and Hessian levels. To formalize this, we extend the River-Valley landscape model by distinguishing between U-shaped valleys (flat) and V-shaped valleys (steep). Based on empirical observations, we conjecture that the recursive architecture of Looped-Attn induces a landscape-level inductive bias towards River-V-Valley. This inductive bias suggest a better loss convergence along the river due to valley hopping, and further encourage learning about complex patterns compared to the River-U-Valley induced by Single-Attn. Building on this insight, we propose SHIFT (Staged HIerarchical Framework for Progressive Training), a principled training strategy that accelerates the training process of Looped-Attn while achieving comparable performances.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10089
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Makes Looped Transformers Perform Better Than Non-Recursive Ones
Gong, Zixuan
Liu, Yong
Teng, Jiaye
Machine Learning
Artificial Intelligence
While looped transformers (termed as Looped-Attn) often outperform standard transformers (termed as Single-Attn) on complex reasoning tasks, the mechanism for this advantage remains underexplored. In this paper, we explain this phenomenon through the lens of loss landscape geometry, inspired by empirical observations of their distinct dynamics at both sample and Hessian levels. To formalize this, we extend the River-Valley landscape model by distinguishing between U-shaped valleys (flat) and V-shaped valleys (steep). Based on empirical observations, we conjecture that the recursive architecture of Looped-Attn induces a landscape-level inductive bias towards River-V-Valley. This inductive bias suggest a better loss convergence along the river due to valley hopping, and further encourage learning about complex patterns compared to the River-U-Valley induced by Single-Attn. Building on this insight, we propose SHIFT (Staged HIerarchical Framework for Progressive Training), a principled training strategy that accelerates the training process of Looped-Attn while achieving comparable performances.
title What Makes Looped Transformers Perform Better Than Non-Recursive Ones
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.10089