Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Haozhe, Xu, Qixin, Liu, Che, Wu, Junhong, Lin, Fangzhen, Chen, Wenhu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915518846337024
author Wang, Haozhe
Xu, Qixin
Liu, Che
Wu, Junhong
Lin, Fangzhen
Chen, Wenhu
author_facet Wang, Haozhe
Xu, Qixin
Liu, Che
Wu, Junhong
Lin, Fangzhen
Chen, Wenhu
contents Reinforcement Learning (RL) has proven highly effective at enhancing the complex reasoning abilities of Large Language Models (LLMs), yet underlying mechanisms driving this success remain largely opaque. Our analysis reveals that puzzling phenomena like ``aha moments", ``length-scaling'' and entropy dynamics are not disparate occurrences but hallmarks of an emergent reasoning hierarchy, akin to the separation of high-level strategic planning from low-level procedural execution in human cognition. We uncover a compelling two-phase dynamic: initially, a model is constrained by procedural correctness and must improve its low-level skills. The learning bottleneck then decisively shifts, with performance gains being driven by the exploration and mastery of high-level strategic planning. This insight exposes a core inefficiency in prevailing RL algorithms like GRPO, which apply optimization pressure agnostically and dilute the learning signal across all tokens. To address this, we propose Hierarchy-Aware Credit Assignment (HICRA), an algorithm that concentrates optimization efforts on high-impact planning tokens. Our extensive experiments validate that HICRA significantly outperforms strong baselines, and offer deep insights into how reasoning advances through the lens of strategic exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03646
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
Wang, Haozhe
Xu, Qixin
Liu, Che
Wu, Junhong
Lin, Fangzhen
Chen, Wenhu
Artificial Intelligence
Computation and Language
Reinforcement Learning (RL) has proven highly effective at enhancing the complex reasoning abilities of Large Language Models (LLMs), yet underlying mechanisms driving this success remain largely opaque. Our analysis reveals that puzzling phenomena like ``aha moments", ``length-scaling'' and entropy dynamics are not disparate occurrences but hallmarks of an emergent reasoning hierarchy, akin to the separation of high-level strategic planning from low-level procedural execution in human cognition. We uncover a compelling two-phase dynamic: initially, a model is constrained by procedural correctness and must improve its low-level skills. The learning bottleneck then decisively shifts, with performance gains being driven by the exploration and mastery of high-level strategic planning. This insight exposes a core inefficiency in prevailing RL algorithms like GRPO, which apply optimization pressure agnostically and dilute the learning signal across all tokens. To address this, we propose Hierarchy-Aware Credit Assignment (HICRA), an algorithm that concentrates optimization efforts on high-impact planning tokens. Our extensive experiments validate that HICRA significantly outperforms strong baselines, and offer deep insights into how reasoning advances through the lens of strategic exploration.
title Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.03646