Language Models Are Implicitly Continuous

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Marro, Samuele, Evangelista, Davide, Huang, X. Angelo, La Malfa, Emanuele, Lombardi, Michele, Wooldridge, Michael
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910904280416256
author Marro, Samuele
Evangelista, Davide
Huang, X. Angelo
La Malfa, Emanuele
Lombardi, Michele
Wooldridge, Michael
author_facet Marro, Samuele
Evangelista, Davide
Huang, X. Angelo
La Malfa, Emanuele
Lombardi, Michele
Wooldridge, Michael
contents Language is typically modelled with discrete sequences. However, the most successful approaches to language modelling, namely neural networks, are continuous and smooth function approximators. In this work, we show that Transformer-based language models implicitly learn to represent sentences as continuous-time functions defined over a continuous input space. This phenomenon occurs in most state-of-the-art Large Language Models (LLMs), including Llama2, Llama3, Phi3, Gemma, Gemma2, and Mistral, and suggests that LLMs reason about language in ways that fundamentally differ from humans. Our work formally extends Transformers to capture the nuances of time and space continuity in both input and output space. Our results challenge the traditional interpretation of how LLMs understand language, with several linguistic and engineering implications.
format Preprint
id arxiv_https___arxiv_org_abs_2504_03933
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Models Are Implicitly Continuous
Marro, Samuele
Evangelista, Davide
Huang, X. Angelo
La Malfa, Emanuele
Lombardi, Michele
Wooldridge, Michael
Computation and Language
Machine Learning
I.2.7; I.2.6
Language is typically modelled with discrete sequences. However, the most successful approaches to language modelling, namely neural networks, are continuous and smooth function approximators. In this work, we show that Transformer-based language models implicitly learn to represent sentences as continuous-time functions defined over a continuous input space. This phenomenon occurs in most state-of-the-art Large Language Models (LLMs), including Llama2, Llama3, Phi3, Gemma, Gemma2, and Mistral, and suggests that LLMs reason about language in ways that fundamentally differ from humans. Our work formally extends Transformers to capture the nuances of time and space continuity in both input and output space. Our results challenge the traditional interpretation of how LLMs understand language, with several linguistic and engineering implications.
title Language Models Are Implicitly Continuous
topic Computation and Language
Machine Learning
I.2.7; I.2.6
url https://arxiv.org/abs/2504.03933