Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Prato, Gabriele, Sodhani, Shagun, Sordoni, Alessandro, Chandar, Sarath
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914204134408192
author Prato, Gabriele
Sodhani, Shagun
Sordoni, Alessandro
Chandar, Sarath
author_facet Prato, Gabriele
Sodhani, Shagun
Sordoni, Alessandro
Chandar, Sarath
contents The standard practice for training large language models involves packing multiple documents together to optimize computational efficiency. However, the impact of this process on the models' capabilities remains largely unexplored. To address this gap, we investigate how different document-packing strategies influence the latent multi-hop reasoning abilities of LLMs. Our findings indicate that packing can improve model performance compared to training on individual documents, at the expense of more compute. To further understand the underlying mechanisms, we conduct an ablation study, identifying key factors that explain the advantages of packing. Ultimately, our research deepens the understanding of LLM training dynamics and provides practical insights for optimizing model development.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14427
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
Prato, Gabriele
Sodhani, Shagun
Sordoni, Alessandro
Chandar, Sarath
Computation and Language
Artificial Intelligence
Machine Learning
The standard practice for training large language models involves packing multiple documents together to optimize computational efficiency. However, the impact of this process on the models' capabilities remains largely unexplored. To address this gap, we investigate how different document-packing strategies influence the latent multi-hop reasoning abilities of LLMs. Our findings indicate that packing can improve model performance compared to training on individual documents, at the expense of more compute. To further understand the underlying mechanisms, we conduct an ablation study, identifying key factors that explain the advantages of packing. Ultimately, our research deepens the understanding of LLM training dynamics and provides practical insights for optimizing model development.
title Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.14427