xVal: A Continuous Numerical Tokenization for Scientific Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Golkar, Siavash, Pettee, Mariel, Eickenberg, Michael, Bietti, Alberto, Cranmer, Miles, Krawezik, Geraud, Lanusse, Francois, McCabe, Michael, Ohana, Ruben, Parker, Liam, Blancard, Bruno Régaldo-Saint, Tesileanu, Tiberiu, Cho, Kyunghyun, Ho, Shirley
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910744067440640
author Golkar, Siavash
Pettee, Mariel
Eickenberg, Michael
Bietti, Alberto
Cranmer, Miles
Krawezik, Geraud
Lanusse, Francois
McCabe, Michael
Ohana, Ruben
Parker, Liam
Blancard, Bruno Régaldo-Saint
Tesileanu, Tiberiu
Cho, Kyunghyun
Ho, Shirley
author_facet Golkar, Siavash
Pettee, Mariel
Eickenberg, Michael
Bietti, Alberto
Cranmer, Miles
Krawezik, Geraud
Lanusse, Francois
McCabe, Michael
Ohana, Ruben
Parker, Liam
Blancard, Bruno Régaldo-Saint
Tesileanu, Tiberiu
Cho, Kyunghyun
Ho, Shirley
contents Due in part to their discontinuous and discrete default encodings for numbers, Large Language Models (LLMs) have not yet been commonly used to process numerically-dense scientific datasets. Rendering datasets as text, however, could help aggregate diverse and multi-modal scientific data into a single training corpus, thereby potentially facilitating the development of foundation models for science. In this work, we introduce xVal, a strategy for continuously tokenizing numbers within language models that results in a more appropriate inductive bias for scientific applications. By training specially-modified language models from scratch on a variety of scientific datasets formatted as text, we find that xVal generally outperforms other common numerical tokenization strategies on metrics including out-of-distribution generalization and computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2310_02989
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle xVal: A Continuous Numerical Tokenization for Scientific Language Models
Golkar, Siavash
Pettee, Mariel
Eickenberg, Michael
Bietti, Alberto
Cranmer, Miles
Krawezik, Geraud
Lanusse, Francois
McCabe, Michael
Ohana, Ruben
Parker, Liam
Blancard, Bruno Régaldo-Saint
Tesileanu, Tiberiu
Cho, Kyunghyun
Ho, Shirley
Machine Learning
Artificial Intelligence
Computation and Language
Due in part to their discontinuous and discrete default encodings for numbers, Large Language Models (LLMs) have not yet been commonly used to process numerically-dense scientific datasets. Rendering datasets as text, however, could help aggregate diverse and multi-modal scientific data into a single training corpus, thereby potentially facilitating the development of foundation models for science. In this work, we introduce xVal, a strategy for continuously tokenizing numbers within language models that results in a more appropriate inductive bias for scientific applications. By training specially-modified language models from scratch on a variety of scientific datasets formatted as text, we find that xVal generally outperforms other common numerical tokenization strategies on metrics including out-of-distribution generalization and computational efficiency.
title xVal: A Continuous Numerical Tokenization for Scientific Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2310.02989