Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Schram, Viktoria, Hiller, Markus, Beck, Daniel, Cohn, Trevor
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2510.16743
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909856469876736
author Schram, Viktoria
Hiller, Markus
Beck, Daniel
Cohn, Trevor
author_facet Schram, Viktoria
Hiller, Markus
Beck, Daniel
Cohn, Trevor
contents The prediction of learning curves for Natural Language Processing (NLP) models enables informed decision-making to meet specific performance objectives, while reducing computational overhead and lowering the costs associated with dataset acquisition and curation. In this work, we formulate the prediction task as a multitask learning problem, where each task's data is modelled as being organized within a two-layer hierarchy. To model the shared information and dependencies across tasks and hierarchical levels, we employ latent variable multi-output Gaussian Processes, enabling to account for task correlations and supporting zero-shot prediction of learning curves (LCs). We demonstrate that this approach facilitates the development of probabilistic scaling laws at lower costs. Applying an active learning strategy, LCs can be queried to reduce predictive uncertainty and provide predictions close to ground truth scaling laws. We validate our framework on three small-scale NLP datasets with up to $30$ LCs. These are obtained from nanoGPT models, from bilingual translation using mBART and Transformer models, and from multilingual translation using M2M100 models of varying sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16743
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot Performance Prediction for Probabilistic Scaling Laws
Schram, Viktoria
Hiller, Markus
Beck, Daniel
Cohn, Trevor
Machine Learning
Computation and Language
The prediction of learning curves for Natural Language Processing (NLP) models enables informed decision-making to meet specific performance objectives, while reducing computational overhead and lowering the costs associated with dataset acquisition and curation. In this work, we formulate the prediction task as a multitask learning problem, where each task's data is modelled as being organized within a two-layer hierarchy. To model the shared information and dependencies across tasks and hierarchical levels, we employ latent variable multi-output Gaussian Processes, enabling to account for task correlations and supporting zero-shot prediction of learning curves (LCs). We demonstrate that this approach facilitates the development of probabilistic scaling laws at lower costs. Applying an active learning strategy, LCs can be queried to reduce predictive uncertainty and provide predictions close to ground truth scaling laws. We validate our framework on three small-scale NLP datasets with up to $30$ LCs. These are obtained from nanoGPT models, from bilingual translation using mBART and Transformer models, and from multilingual translation using M2M100 models of varying sizes.
title Zero-Shot Performance Prediction for Probabilistic Scaling Laws
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2510.16743