Language models can learn implicit multi-hop reasoning, but only if they have lots of training data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Yuekun, Du, Yupei, Zhu, Dawei, Hahn, Michael, Koller, Alexander
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918321731928064
author Yao, Yuekun
Du, Yupei
Zhu, Dawei
Hahn, Michael
Koller, Alexander
author_facet Yao, Yuekun
Du, Yupei
Zhu, Dawei
Hahn, Michael
Koller, Alexander
contents Implicit reasoning is the ability of a language model to solve multi-hop reasoning tasks in a single forward pass, without chain of thought. We investigate this capability using GPT2-style language models trained from scratch on controlled $k$-hop reasoning datasets ($k = 2, 3, 4$). We show that while such models can indeed learn implicit $k$-hop reasoning, the required training data grows exponentially in $k$, and the required number of transformer layers grows linearly in $k$. We offer a theoretical explanation for why this depth growth is necessary. We further find that the data requirement can be mitigated, but not eliminated, through curriculum learning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17923
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language models can learn implicit multi-hop reasoning, but only if they have lots of training data
Yao, Yuekun
Du, Yupei
Zhu, Dawei
Hahn, Michael
Koller, Alexander
Computation and Language
Implicit reasoning is the ability of a language model to solve multi-hop reasoning tasks in a single forward pass, without chain of thought. We investigate this capability using GPT2-style language models trained from scratch on controlled $k$-hop reasoning datasets ($k = 2, 3, 4$). We show that while such models can indeed learn implicit $k$-hop reasoning, the required training data grows exponentially in $k$, and the required number of transformer layers grows linearly in $k$. We offer a theoretical explanation for why this depth growth is necessary. We further find that the data requirement can be mitigated, but not eliminated, through curriculum learning.
title Language models can learn implicit multi-hop reasoning, but only if they have lots of training data
topic Computation and Language
url https://arxiv.org/abs/2505.17923