Auxiliary task demands mask the capabilities of smaller language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Jennifer, Frank, Michael C.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914891680448512
author Hu, Jennifer
Frank, Michael C.
author_facet Hu, Jennifer
Frank, Michael C.
contents Developmental psychologists have argued about when cognitive capacities such as language understanding or theory of mind emerge. These debates often hinge on the concept of "task demands" -- the auxiliary challenges associated with performing a particular evaluation -- that may mask the child's underlying ability. The same issues arise when measuring the capacities of language models (LMs): performance on a task is a function of the model's underlying knowledge, combined with the model's ability to interpret and perform the task given its available resources. Here, we show that for analogical reasoning, reflective reasoning, word prediction, and grammaticality judgments, evaluation methods with greater task demands yield lower performance than evaluations with reduced demands. This "demand gap" is most pronounced for models with fewer parameters and less training data. Our results illustrate that LM performance should not be interpreted as a direct indication of intelligence (or lack thereof), but as a reflection of capacities seen through the lens of researchers' design choices.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02418
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Auxiliary task demands mask the capabilities of smaller language models
Hu, Jennifer
Frank, Michael C.
Computation and Language
Artificial Intelligence
Developmental psychologists have argued about when cognitive capacities such as language understanding or theory of mind emerge. These debates often hinge on the concept of "task demands" -- the auxiliary challenges associated with performing a particular evaluation -- that may mask the child's underlying ability. The same issues arise when measuring the capacities of language models (LMs): performance on a task is a function of the model's underlying knowledge, combined with the model's ability to interpret and perform the task given its available resources. Here, we show that for analogical reasoning, reflective reasoning, word prediction, and grammaticality judgments, evaluation methods with greater task demands yield lower performance than evaluations with reduced demands. This "demand gap" is most pronounced for models with fewer parameters and less training data. Our results illustrate that LM performance should not be interpreted as a direct indication of intelligence (or lack thereof), but as a reflection of capacities seen through the lens of researchers' design choices.
title Auxiliary task demands mask the capabilities of smaller language models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.02418