Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kapl, Ferdinand, Angelis, Emmanouil, Höppe, Tobias, Maile, Kaitlin, von Oswald, Johannes, Scherrer, Nino, Bauer, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915664002809856
author Kapl, Ferdinand
Angelis, Emmanouil
Höppe, Tobias
Maile, Kaitlin
von Oswald, Johannes
Scherrer, Nino
Bauer, Stefan
author_facet Kapl, Ferdinand
Angelis, Emmanouil
Höppe, Tobias
Maile, Kaitlin
von Oswald, Johannes
Scherrer, Nino
Bauer, Stefan
contents Gradually growing the depth of Transformers during training can not only reduce training cost but also lead to improved reasoning performance, as shown by MIDAS (Saunshi et al., 2024). Thus far, however, a mechanistic understanding of these gains has been missing. In this work, we establish a connection to recent work showing that layers in the second half of non-grown, pre-layernorm Transformers contribute much less to the final output distribution than those in the first half - also known as the Curse of Depth (Sun et al., 2025, Csordás et al., 2025). Using depth-wise analyses, we demonstrate that growth via gradual middle stacking yields more effective utilization of model depth, alters the residual stream structure, and facilitates the formation of permutable computational blocks. In addition, we propose a lightweight modification of MIDAS that yields further improvements in downstream reasoning benchmarks. Overall, this work highlights how the gradual growth of model depth can lead to the formation of distinct computational circuits and overcome the limited depth utilization seen in standard non-grown models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08819
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
Kapl, Ferdinand
Angelis, Emmanouil
Höppe, Tobias
Maile, Kaitlin
von Oswald, Johannes
Scherrer, Nino
Bauer, Stefan
Computation and Language
Artificial Intelligence
Machine Learning
Gradually growing the depth of Transformers during training can not only reduce training cost but also lead to improved reasoning performance, as shown by MIDAS (Saunshi et al., 2024). Thus far, however, a mechanistic understanding of these gains has been missing. In this work, we establish a connection to recent work showing that layers in the second half of non-grown, pre-layernorm Transformers contribute much less to the final output distribution than those in the first half - also known as the Curse of Depth (Sun et al., 2025, Csordás et al., 2025). Using depth-wise analyses, we demonstrate that growth via gradual middle stacking yields more effective utilization of model depth, alters the residual stream structure, and facilitates the formation of permutable computational blocks. In addition, we propose a lightweight modification of MIDAS that yields further improvements in downstream reasoning benchmarks. Overall, this work highlights how the gradual growth of model depth can lead to the formation of distinct computational circuits and overcome the limited depth utilization seen in standard non-grown models.
title Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.08819