Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Kapl, Ferdinand, Angelis, Emmanouil, Höppe, Tobias, Maile, Kaitlin, von Oswald, Johannes, Scherrer, Nino, Bauer, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Growing to Looping: A Unified View of Iterative Computation in LLMs
by: Kapl, Ferdinand, et al.
Published: (2026)
by: Kapl, Ferdinand, et al.
Published: (2026)
From Words to Amino Acids: Does the Curse of Depth Persist?
by: Siji, Aleena, et al.
Published: (2026)
by: Siji, Aleena, et al.
Published: (2026)
How Do LLMs Use Their Depth?
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
The Curse of Depth in Large Language Models
by: Sun, Wenfang, et al.
Published: (2025)
by: Sun, Wenfang, et al.
Published: (2025)
Mixture-of-Depths Attention
by: Zhu, Lianghui, et al.
Published: (2026)
by: Zhu, Lianghui, et al.
Published: (2026)
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
by: von Oswald, Johannes, et al.
Published: (2025)
by: von Oswald, Johannes, et al.
Published: (2025)
Tuning Language Models by Mixture-of-Depths Ensemble
by: Luo, Haoyan, et al.
Published: (2024)
by: Luo, Haoyan, et al.
Published: (2024)
Prompt-based Depth Pruning of Large Language Models
by: Wee, Juyun, et al.
Published: (2025)
by: Wee, Juyun, et al.
Published: (2025)
DepthCharge: A Domain-Agnostic Framework for Measuring Depth-Dependent Knowledge in Large Language Models
by: Sheppert, Alexander
Published: (2026)
by: Sheppert, Alexander
Published: (2026)
Single-Pass, Depth-Selective Reading for Multi-Aspect Sentiment Analysis
by: Xia, Yan, et al.
Published: (2026)
by: Xia, Yan, et al.
Published: (2026)
Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation
by: Chen, Yangneng, et al.
Published: (2026)
by: Chen, Yangneng, et al.
Published: (2026)
DND: Boosting Large Language Models with Dynamic Nested Depth
by: Chen, Tieyuan, et al.
Published: (2025)
by: Chen, Tieyuan, et al.
Published: (2025)
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)
by: Lopardo, Gianluigi, et al.
Published: (2022)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
by: Brown, Oscar, et al.
Published: (2024)
by: Brown, Oscar, et al.
Published: (2024)
WorDepth: Variational Language Prior for Monocular Depth Estimation
by: Zeng, Ziyao, et al.
Published: (2024)
by: Zeng, Ziyao, et al.
Published: (2024)
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
by: Wu, Zongqian, et al.
Published: (2025)
by: Wu, Zongqian, et al.
Published: (2025)
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
by: Ohmer, Xenia, et al.
Published: (2024)
by: Ohmer, Xenia, et al.
Published: (2024)
Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus
by: Lee, Seungpil, et al.
Published: (2024)
by: Lee, Seungpil, et al.
Published: (2024)
Layers at Similar Depths Generate Similar Activations Across LLM Architectures
by: Wolfram, Christopher, et al.
Published: (2025)
by: Wolfram, Christopher, et al.
Published: (2025)
The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning
by: Xu, Yi, et al.
Published: (2026)
by: Xu, Yi, et al.
Published: (2026)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
by: Lu, Wenquan, et al.
Published: (2025)
by: Lu, Wenquan, et al.
Published: (2025)
Measuring the Depth of LLM Unlearning via Activation Patching
by: Lee, Jaeung, et al.
Published: (2026)
by: Lee, Jaeung, et al.
Published: (2026)
Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents
by: Yang, Ruihan, et al.
Published: (2026)
by: Yang, Ruihan, et al.
Published: (2026)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
by: Liu, Akide, et al.
Published: (2024)
by: Liu, Akide, et al.
Published: (2024)
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
Depth-Width tradeoffs in Algorithmic Reasoning of Graph Tasks with Transformers
by: Yehudai, Gilad, et al.
Published: (2025)
by: Yehudai, Gilad, et al.
Published: (2025)
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
by: Chen, Hung-Hsuan
Published: (2026)
by: Chen, Hung-Hsuan
Published: (2026)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
by: Kohli, Harsh, et al.
Published: (2026)
by: Kohli, Harsh, et al.
Published: (2026)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
by: Dehghanighobadi, Zahra, et al.
Published: (2026)
by: Dehghanighobadi, Zahra, et al.
Published: (2026)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
by: Knupp, Jonas, et al.
Published: (2026)
by: Knupp, Jonas, et al.
Published: (2026)
GRAPHMOE: Amplifying Cognitive Depth of Mixture-of-Experts Network via Introducing Self-Rethinking Mechanism
by: Lv, Bo, et al.
Published: (2025)
by: Lv, Bo, et al.
Published: (2025)
Depth $F_1$: Improving Evaluation of Cross-Domain Text Classification by Measuring Semantic Generalizability
by: Seegmiller, Parker, et al.
Published: (2024)
by: Seegmiller, Parker, et al.
Published: (2024)
When Does Sparsity Mitigate the Curse of Depth in LLMs
by: Muhtar, Dilxat, et al.
Published: (2026)
by: Muhtar, Dilxat, et al.
Published: (2026)
Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers?
by: Jin, Mingyu, et al.
Published: (2024)
by: Jin, Mingyu, et al.
Published: (2024)
An Analysis and Mitigation of the Reversal Curse
by: Lv, Ang, et al.
Published: (2023)
by: Lv, Ang, et al.
Published: (2023)
Do Language Models Use Their Depth Efficiently?
by: Csordás, Róbert, et al.
Published: (2025)
by: Csordás, Róbert, et al.
Published: (2025)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
Dynamic Reasoning Chains through Depth-Specialized Mixture-of-Experts in Transformer Architectures
by: Roy, Sampurna, et al.
Published: (2025)
by: Roy, Sampurna, et al.
Published: (2025)
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
by: Shopkhoev, Dmitriy, et al.
Published: (2025)
by: Shopkhoev, Dmitriy, et al.
Published: (2025)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
by: Chen, Yilong, et al.
Published: (2026)
by: Chen, Yilong, et al.
Published: (2026)
Similar Items
-
From Growing to Looping: A Unified View of Iterative Computation in LLMs
by: Kapl, Ferdinand, et al.
Published: (2026) -
From Words to Amino Acids: Does the Curse of Depth Persist?
by: Siji, Aleena, et al.
Published: (2026) -
How Do LLMs Use Their Depth?
by: Gupta, Akshat, et al.
Published: (2025) -
The Curse of Depth in Large Language Models
by: Sun, Wenfang, et al.
Published: (2025) -
Mixture-of-Depths Attention
by: Zhu, Lianghui, et al.
Published: (2026)