Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruis, Laura, Mozes, Maximilian, Bae, Juhan, Kamalakara, Siddhartha Rao, Talupuru, Dwarak, Locatelli, Acyr, Kirk, Robert, Rocktäschel, Tim, Grefenstette, Edward, Bartolo, Max
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917946796802048
author Ruis, Laura
Mozes, Maximilian
Bae, Juhan
Kamalakara, Siddhartha Rao
Talupuru, Dwarak
Locatelli, Acyr
Kirk, Robert
Rocktäschel, Tim
Grefenstette, Edward
Bartolo, Max
author_facet Ruis, Laura
Mozes, Maximilian
Bae, Juhan
Kamalakara, Siddhartha Rao
Talupuru, Dwarak
Locatelli, Acyr
Kirk, Robert
Rocktäschel, Tim
Grefenstette, Edward
Bartolo, Max
contents The capabilities and limitations of Large Language Models have been sketched out in great detail in recent years, providing an intriguing yet conflicting picture. On the one hand, LLMs demonstrate a general ability to solve problems. On the other hand, they show surprising reasoning gaps when compared to humans, casting doubt on the robustness of their generalisation strategies. The sheer volume of data used in the design of LLMs has precluded us from applying the method traditionally used to measure generalisation: train-test set separation. To overcome this, we study what kind of generalisation strategies LLMs employ when performing reasoning tasks by investigating the pretraining data they rely on. For two models of different sizes (7B and 35B) and 2.5B of their pretraining tokens, we identify what documents influence the model outputs for three simple mathematical reasoning tasks and contrast this to the data that are influential for answering factual questions. We find that, while the models rely on mostly distinct sets of data for each factual question, a document often has a similar influence across different reasoning questions within the same task, indicating the presence of procedural knowledge. We further find that the answers to factual questions often show up in the most influential data. However, for reasoning questions the answers usually do not show up as highly influential, nor do the answers to the intermediate reasoning steps. When we characterise the top ranked documents for the reasoning questions qualitatively, we confirm that the influential documents often contain procedural knowledge, like demonstrating how to obtain a solution using formulae or code. Our findings indicate that the approach to reasoning the models use is unlike retrieval, and more like a generalisable strategy that synthesises procedural knowledge from documents doing a similar form of reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12580
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
Ruis, Laura
Mozes, Maximilian
Bae, Juhan
Kamalakara, Siddhartha Rao
Talupuru, Dwarak
Locatelli, Acyr
Kirk, Robert
Rocktäschel, Tim
Grefenstette, Edward
Bartolo, Max
Computation and Language
Machine Learning
The capabilities and limitations of Large Language Models have been sketched out in great detail in recent years, providing an intriguing yet conflicting picture. On the one hand, LLMs demonstrate a general ability to solve problems. On the other hand, they show surprising reasoning gaps when compared to humans, casting doubt on the robustness of their generalisation strategies. The sheer volume of data used in the design of LLMs has precluded us from applying the method traditionally used to measure generalisation: train-test set separation. To overcome this, we study what kind of generalisation strategies LLMs employ when performing reasoning tasks by investigating the pretraining data they rely on. For two models of different sizes (7B and 35B) and 2.5B of their pretraining tokens, we identify what documents influence the model outputs for three simple mathematical reasoning tasks and contrast this to the data that are influential for answering factual questions. We find that, while the models rely on mostly distinct sets of data for each factual question, a document often has a similar influence across different reasoning questions within the same task, indicating the presence of procedural knowledge. We further find that the answers to factual questions often show up in the most influential data. However, for reasoning questions the answers usually do not show up as highly influential, nor do the answers to the intermediate reasoning steps. When we characterise the top ranked documents for the reasoning questions qualitatively, we confirm that the influential documents often contain procedural knowledge, like demonstrating how to obtain a solution using formulae or code. Our findings indicate that the approach to reasoning the models use is unlike retrieval, and more like a generalisable strategy that synthesises procedural knowledge from documents doing a similar form of reasoning.
title Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.12580