xVal: A Continuous Numerical Tokenization for Scientific Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910744067440640 |
|---|---|
| author | Golkar, Siavash Pettee, Mariel Eickenberg, Michael Bietti, Alberto Cranmer, Miles Krawezik, Geraud Lanusse, Francois McCabe, Michael Ohana, Ruben Parker, Liam Blancard, Bruno Régaldo-Saint Tesileanu, Tiberiu Cho, Kyunghyun Ho, Shirley |
| author_facet | Golkar, Siavash Pettee, Mariel Eickenberg, Michael Bietti, Alberto Cranmer, Miles Krawezik, Geraud Lanusse, Francois McCabe, Michael Ohana, Ruben Parker, Liam Blancard, Bruno Régaldo-Saint Tesileanu, Tiberiu Cho, Kyunghyun Ho, Shirley |
| contents | Due in part to their discontinuous and discrete default encodings for numbers, Large Language Models (LLMs) have not yet been commonly used to process numerically-dense scientific datasets. Rendering datasets as text, however, could help aggregate diverse and multi-modal scientific data into a single training corpus, thereby potentially facilitating the development of foundation models for science. In this work, we introduce xVal, a strategy for continuously tokenizing numbers within language models that results in a more appropriate inductive bias for scientific applications. By training specially-modified language models from scratch on a variety of scientific datasets formatted as text, we find that xVal generally outperforms other common numerical tokenization strategies on metrics including out-of-distribution generalization and computational efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_02989 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | xVal: A Continuous Numerical Tokenization for Scientific Language Models Golkar, Siavash Pettee, Mariel Eickenberg, Michael Bietti, Alberto Cranmer, Miles Krawezik, Geraud Lanusse, Francois McCabe, Michael Ohana, Ruben Parker, Liam Blancard, Bruno Régaldo-Saint Tesileanu, Tiberiu Cho, Kyunghyun Ho, Shirley Machine Learning Artificial Intelligence Computation and Language Due in part to their discontinuous and discrete default encodings for numbers, Large Language Models (LLMs) have not yet been commonly used to process numerically-dense scientific datasets. Rendering datasets as text, however, could help aggregate diverse and multi-modal scientific data into a single training corpus, thereby potentially facilitating the development of foundation models for science. In this work, we introduce xVal, a strategy for continuously tokenizing numbers within language models that results in a more appropriate inductive bias for scientific applications. By training specially-modified language models from scratch on a variety of scientific datasets formatted as text, we find that xVal generally outperforms other common numerical tokenization strategies on metrics including out-of-distribution generalization and computational efficiency. |
| title | xVal: A Continuous Numerical Tokenization for Scientific Language Models |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2310.02989 |