Exploring Internal Numeracy in Language Models: A Case Study on ALBERT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wennberg, Ulme, Henter, Gustav Eje
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910423161241600
author Wennberg, Ulme
Henter, Gustav Eje
author_facet Wennberg, Ulme
Henter, Gustav Eje
contents It has been found that Transformer-based language models have the ability to perform basic quantitative reasoning. In this paper, we propose a method for studying how these models internally represent numerical data, and use our proposal to analyze the ALBERT family of language models. Specifically, we extract the learned embeddings these models use to represent tokens that correspond to numbers and ordinals, and subject these embeddings to Principal Component Analysis (PCA). PCA results reveal that ALBERT models of different sizes, trained and initialized separately, consistently learn to use the axes of greatest variation to represent the approximate ordering of various numerical concepts. Numerals and their textual counterparts are represented in separate clusters, but increase along the same direction in 2D space. Our findings illustrate that language models, trained purely to model text, can intuit basic mathematical concepts, opening avenues for NLP applications that intersect with quantitative reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16574
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Internal Numeracy in Language Models: A Case Study on ALBERT
Wennberg, Ulme
Henter, Gustav Eje
Computation and Language
It has been found that Transformer-based language models have the ability to perform basic quantitative reasoning. In this paper, we propose a method for studying how these models internally represent numerical data, and use our proposal to analyze the ALBERT family of language models. Specifically, we extract the learned embeddings these models use to represent tokens that correspond to numbers and ordinals, and subject these embeddings to Principal Component Analysis (PCA). PCA results reveal that ALBERT models of different sizes, trained and initialized separately, consistently learn to use the axes of greatest variation to represent the approximate ordering of various numerical concepts. Numerals and their textual counterparts are represented in separate clusters, but increase along the same direction in 2D space. Our findings illustrate that language models, trained purely to model text, can intuit basic mathematical concepts, opening avenues for NLP applications that intersect with quantitative reasoning.
title Exploring Internal Numeracy in Language Models: A Case Study on ALBERT
topic Computation and Language
url https://arxiv.org/abs/2404.16574