Interleaving Text and Number Embeddings to Solve Mathemathics Problems

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Alberts, Marvin, Gabrieli, Gianmarco, Morales, Irina Espejo
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914988698894336
author Alberts, Marvin
Gabrieli, Gianmarco
Morales, Irina Espejo
author_facet Alberts, Marvin
Gabrieli, Gianmarco
Morales, Irina Espejo
contents Integrating text and numbers effectively is a crucial step towards enhancing Large Language Models (LLMs) capabilities in assisting in scientific tasks. While most current approaches rely on discrete tokenization of numbers, for instance, conversion to scientific notation or base 10-decomposition, a recent approach proposed a continuous numerical encoding as an inductive bias. In this paper, we build upon this approach by introducing more expressive numerical embeddings. Our method addresses key shortcomings, including the elimination of numerical artefacts and the ability to handle a wide range of magnitudes without clipping. Our work presents two key contributions. First, we employ an MLP to assign distinct directions in the embedding space to different numbers. Our second contribution is the introduction of a routing layer that differentiates between numerical and text embeddings. We hypothesise that this combined approach enables the model to distinguish between text and number distributions while maintaining its capacity for arithmetic operations. Using only a 45 M parameter encoder-decoder architecture our method achieves a $R^2$=0.9988 over a wide range of magnitude ($10^{-3},10^{8}$). In addition, we empirically observe a reduction of the numerical artefacts and biases observed compared to the baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19353
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Interleaving Text and Number Embeddings to Solve Mathemathics Problems
Alberts, Marvin
Gabrieli, Gianmarco
Morales, Irina Espejo
Computation and Language
Artificial Intelligence
Integrating text and numbers effectively is a crucial step towards enhancing Large Language Models (LLMs) capabilities in assisting in scientific tasks. While most current approaches rely on discrete tokenization of numbers, for instance, conversion to scientific notation or base 10-decomposition, a recent approach proposed a continuous numerical encoding as an inductive bias. In this paper, we build upon this approach by introducing more expressive numerical embeddings. Our method addresses key shortcomings, including the elimination of numerical artefacts and the ability to handle a wide range of magnitudes without clipping. Our work presents two key contributions. First, we employ an MLP to assign distinct directions in the embedding space to different numbers. Our second contribution is the introduction of a routing layer that differentiates between numerical and text embeddings. We hypothesise that this combined approach enables the model to distinguish between text and number distributions while maintaining its capacity for arithmetic operations. Using only a 45 M parameter encoder-decoder architecture our method achieves a $R^2$=0.9988 over a wide range of magnitude ($10^{-3},10^{8}$). In addition, we empirically observe a reduction of the numerical artefacts and biases observed compared to the baselines.
title Interleaving Text and Number Embeddings to Solve Mathemathics Problems
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.19353