Form and Meaning in Intrinsic Multilingual Evaluations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Poelman, Wessel, de Lhoneux, Miryam
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917205682159616
author Poelman, Wessel
de Lhoneux, Miryam
author_facet Poelman, Wessel
de Lhoneux, Miryam
contents Intrinsic evaluation metrics for conditional language models, such as perplexity or bits-per-character, are widely used in both mono- and multilingual settings. These metrics are rather straightforward to use and compare in monolingual setups, but rest on a number of assumptions in multilingual setups. One such assumption is that comparing the perplexity of CLMs on parallel sentences is indicative of their quality since the information content (here understood as the semantic meaning) is the same. However, the metrics are inherently measuring information content in the information-theoretic sense. We make this and other such assumptions explicit and discuss their implications. We perform experiments with six metrics on two multi-parallel corpora both with mono- and multilingual models. Ultimately, we find that current metrics are not universally comparable. We look at the form-meaning debate to provide some explanation for this.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10580
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Form and Meaning in Intrinsic Multilingual Evaluations
Poelman, Wessel
de Lhoneux, Miryam
Computation and Language
Intrinsic evaluation metrics for conditional language models, such as perplexity or bits-per-character, are widely used in both mono- and multilingual settings. These metrics are rather straightforward to use and compare in monolingual setups, but rest on a number of assumptions in multilingual setups. One such assumption is that comparing the perplexity of CLMs on parallel sentences is indicative of their quality since the information content (here understood as the semantic meaning) is the same. However, the metrics are inherently measuring information content in the information-theoretic sense. We make this and other such assumptions explicit and discuss their implications. We perform experiments with six metrics on two multi-parallel corpora both with mono- and multilingual models. Ultimately, we find that current metrics are not universally comparable. We look at the form-meaning debate to provide some explanation for this.
title Form and Meaning in Intrinsic Multilingual Evaluations
topic Computation and Language
url https://arxiv.org/abs/2601.10580