Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914204134408192 |
|---|---|
| author | Prato, Gabriele Sodhani, Shagun Sordoni, Alessandro Chandar, Sarath |
| author_facet | Prato, Gabriele Sodhani, Shagun Sordoni, Alessandro Chandar, Sarath |
| contents | The standard practice for training large language models involves packing multiple documents together to optimize computational efficiency. However, the impact of this process on the models' capabilities remains largely unexplored. To address this gap, we investigate how different document-packing strategies influence the latent multi-hop reasoning abilities of LLMs. Our findings indicate that packing can improve model performance compared to training on individual documents, at the expense of more compute. To further understand the underlying mechanisms, we conduct an ablation study, identifying key factors that explain the advantages of packing. Ultimately, our research deepens the understanding of LLM training dynamics and provides practical insights for optimizing model development. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_14427 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models Prato, Gabriele Sodhani, Shagun Sordoni, Alessandro Chandar, Sarath Computation and Language Artificial Intelligence Machine Learning The standard practice for training large language models involves packing multiple documents together to optimize computational efficiency. However, the impact of this process on the models' capabilities remains largely unexplored. To address this gap, we investigate how different document-packing strategies influence the latent multi-hop reasoning abilities of LLMs. Our findings indicate that packing can improve model performance compared to training on individual documents, at the expense of more compute. To further understand the underlying mechanisms, we conduct an ablation study, identifying key factors that explain the advantages of packing. Ultimately, our research deepens the understanding of LLM training dynamics and provides practical insights for optimizing model development. |
| title | Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2512.14427 |