Do Generalisation Results Generalise?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boglioni, Matteo, Sgobbi, Andrea, Tavernini, Gabriel, Rita, Francesco, Mosbach, Marius, Pimentel, Tiago
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912754451873792
author Boglioni, Matteo
Sgobbi, Andrea
Tavernini, Gabriel
Rita, Francesco
Mosbach, Marius
Pimentel, Tiago
author_facet Boglioni, Matteo
Sgobbi, Andrea
Tavernini, Gabriel
Rita, Francesco
Mosbach, Marius
Pimentel, Tiago
contents A large language model's (LLM's) out-of-distribution (OOD) generalisation ability is crucial to its deployment. Previous work assessing LLMs' generalisation performance, however, typically focuses on a single out-of-distribution dataset. This approach may fail to precisely evaluate the capabilities of the model, as the data shifts encountered once a model is deployed are much more diverse. In this work, we investigate whether OOD generalisation results generalise. More specifically, we evaluate a model's performance across multiple OOD testsets throughout a finetuning run; we then evaluate the partial correlation of performances across these testsets, regressing out in-domain performance. This allows us to assess how correlated are generalisation performances once in-domain performance is controlled for. Analysing OLMo2 and OPT, we observe no overarching trend in generalisation results: the existence of a positive or negative correlation between any two OOD testsets depends strongly on the specific choice of model analysed.
format Preprint
id arxiv_https___arxiv_org_abs_2512_07832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Generalisation Results Generalise?
Boglioni, Matteo
Sgobbi, Andrea
Tavernini, Gabriel
Rita, Francesco
Mosbach, Marius
Pimentel, Tiago
Computation and Language
Machine Learning
A large language model's (LLM's) out-of-distribution (OOD) generalisation ability is crucial to its deployment. Previous work assessing LLMs' generalisation performance, however, typically focuses on a single out-of-distribution dataset. This approach may fail to precisely evaluate the capabilities of the model, as the data shifts encountered once a model is deployed are much more diverse. In this work, we investigate whether OOD generalisation results generalise. More specifically, we evaluate a model's performance across multiple OOD testsets throughout a finetuning run; we then evaluate the partial correlation of performances across these testsets, regressing out in-domain performance. This allows us to assess how correlated are generalisation performances once in-domain performance is controlled for. Analysing OLMo2 and OPT, we observe no overarching trend in generalisation results: the existence of a positive or negative correlation between any two OOD testsets depends strongly on the specific choice of model analysed.
title Do Generalisation Results Generalise?
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2512.07832