On Characterizations for Language Generation: Interplay of Hallucinations, Breadth, and Stability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kalavasis, Alkis, Mehrotra, Anay, Velegkas, Grigoris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908430720040960
author Kalavasis, Alkis
Mehrotra, Anay
Velegkas, Grigoris
author_facet Kalavasis, Alkis
Mehrotra, Anay
Velegkas, Grigoris
contents We study language generation in the limit - introduced by Kleinberg and Mullainathan [KM24] - building on classical works of Gold [Gol67] and Angluin [Ang79]. [KM24]'s main result is an algorithm for generating from any countable language collection in the limit. While their algorithm eventually generates unseen strings from the target language $K$, it sacrifices coverage or breadth, i.e., its ability to generate a rich set of strings. Recent work introduces different notions of breadth and explores when generation with breadth is possible, leaving a full characterization of these notions open. Our first set of results settles this by characterizing generation for existing notions of breadth and their natural extensions. Interestingly, our lower bounds are very flexible and hold for many performance metrics beyond breadth - for instance, showing that, in general, it is impossible to train generators which achieve a higher perplexity or lower hallucination rate for $K$ compared to other languages. Next, we study language generation with breadth and stable generators - algorithms that eventually stop changing after seeing an arbitrary but finite number of strings - and prove unconditional lower bounds for such generators, strengthening the results of [KMV25] and demonstrating that generation with many existing notions of breadth becomes equally hard, when stability is required. This gives a separation for generation with approximate breadth, between stable and unstable generators, highlighting the rich interplay between breadth, stability, and consistency in language generation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18530
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Characterizations for Language Generation: Interplay of Hallucinations, Breadth, and Stability
Kalavasis, Alkis
Mehrotra, Anay
Velegkas, Grigoris
Machine Learning
Artificial Intelligence
Computation and Language
Data Structures and Algorithms
We study language generation in the limit - introduced by Kleinberg and Mullainathan [KM24] - building on classical works of Gold [Gol67] and Angluin [Ang79]. [KM24]'s main result is an algorithm for generating from any countable language collection in the limit. While their algorithm eventually generates unseen strings from the target language $K$, it sacrifices coverage or breadth, i.e., its ability to generate a rich set of strings. Recent work introduces different notions of breadth and explores when generation with breadth is possible, leaving a full characterization of these notions open. Our first set of results settles this by characterizing generation for existing notions of breadth and their natural extensions. Interestingly, our lower bounds are very flexible and hold for many performance metrics beyond breadth - for instance, showing that, in general, it is impossible to train generators which achieve a higher perplexity or lower hallucination rate for $K$ compared to other languages. Next, we study language generation with breadth and stable generators - algorithms that eventually stop changing after seeing an arbitrary but finite number of strings - and prove unconditional lower bounds for such generators, strengthening the results of [KMV25] and demonstrating that generation with many existing notions of breadth becomes equally hard, when stability is required. This gives a separation for generation with approximate breadth, between stable and unstable generators, highlighting the rich interplay between breadth, stability, and consistency in language generation.
title On Characterizations for Language Generation: Interplay of Hallucinations, Breadth, and Stability
topic Machine Learning
Artificial Intelligence
Computation and Language
Data Structures and Algorithms
url https://arxiv.org/abs/2412.18530