DND: Boosting Large Language Models with Dynamic Nested Depth

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Tieyuan, Chen, Xiaodong, Chen, Haoxing, Lan, Zhenzhong, Lin, Weiyao, Li, Jianguo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911400903835648
author Chen, Tieyuan
Chen, Xiaodong
Chen, Haoxing
Lan, Zhenzhong
Lin, Weiyao
Li, Jianguo
author_facet Chen, Tieyuan
Chen, Xiaodong
Chen, Haoxing
Lan, Zhenzhong
Lin, Weiyao
Li, Jianguo
contents We introduce Dynamic Nested Depth (DND), a novel method that improves performance for off-the-shelf LLMs by selecting critical tokens to reprocess in a nested depth manner. Specifically, at the end of the given transformer layer, DND identifies more critical tokens with a router and feeds them back for an extra round of processing, effectively ``reviewing" difficult tokens while avoiding redundant computation for easier ones. The dynamic selection mechanism is tailored for precise control via two novel strategies: a router controlling loss to enhance token selection distinguishability, and a threshold control scheme to ensure selection stability. We demonstrate the effectiveness of DND by directly integrating it into pre-trained dense and MoE models during a post-training phase. On diverse benchmarks, this approach boosts the performances of the dense Qwen3-1.7B by 1.88% and the MoE Qwen3-30B-A3B by 0.87%, all with a minimal parameter and computing increase.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DND: Boosting Large Language Models with Dynamic Nested Depth
Chen, Tieyuan
Chen, Xiaodong
Chen, Haoxing
Lan, Zhenzhong
Lin, Weiyao
Li, Jianguo
Computation and Language
Artificial Intelligence
We introduce Dynamic Nested Depth (DND), a novel method that improves performance for off-the-shelf LLMs by selecting critical tokens to reprocess in a nested depth manner. Specifically, at the end of the given transformer layer, DND identifies more critical tokens with a router and feeds them back for an extra round of processing, effectively ``reviewing" difficult tokens while avoiding redundant computation for easier ones. The dynamic selection mechanism is tailored for precise control via two novel strategies: a router controlling loss to enhance token selection distinguishability, and a threshold control scheme to ensure selection stability. We demonstrate the effectiveness of DND by directly integrating it into pre-trained dense and MoE models during a post-training phase. On diverse benchmarks, this approach boosts the performances of the dense Qwen3-1.7B by 1.88% and the MoE Qwen3-30B-A3B by 0.87%, all with a minimal parameter and computing increase.
title DND: Boosting Large Language Models with Dynamic Nested Depth
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.11001