Predicting Language Models' Success at Zero-Shot Probabilistic Prediction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ren, Kevin, Cortes-Gomez, Santiago, Patiño, Carlos Miguel, Joshi, Ananya, Lyu, Ruiqi, Tang, Jingjing, Turcan, Alistair, Yamin, Khurram, Wu, Steven, Wilder, Bryan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911163230453760
author Ren, Kevin
Cortes-Gomez, Santiago
Patiño, Carlos Miguel
Joshi, Ananya
Lyu, Ruiqi
Tang, Jingjing
Turcan, Alistair
Yamin, Khurram
Wu, Steven
Wilder, Bryan
author_facet Ren, Kevin
Cortes-Gomez, Santiago
Patiño, Carlos Miguel
Joshi, Ananya
Lyu, Ruiqi
Tang, Jingjing
Turcan, Alistair
Yamin, Khurram
Wu, Steven
Wilder, Bryan
contents Recent work has investigated the capabilities of large language models (LLMs) as zero-shot models for generating individual-level characteristics (e.g., to serve as risk models or augment survey datasets). However, when should a user have confidence that an LLM will provide high-quality predictions for their particular task? To address this question, we conduct a large-scale empirical study of LLMs' zero-shot predictive capabilities across a wide range of tabular prediction tasks. We find that LLMs' performance is highly variable, both on tasks within the same dataset and across different datasets. However, when the LLM performs well on the base prediction task, its predicted probabilities become a stronger signal for individual-level accuracy. Then, we construct metrics to predict LLMs' performance at the task level, aiming to distinguish between tasks where LLMs may perform well and where they are likely unsuitable. We find that some of these metrics, each of which are assessed without labeled data, yield strong signals of LLMs' predictive performance on new tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15356
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Predicting Language Models' Success at Zero-Shot Probabilistic Prediction
Ren, Kevin
Cortes-Gomez, Santiago
Patiño, Carlos Miguel
Joshi, Ananya
Lyu, Ruiqi
Tang, Jingjing
Turcan, Alistair
Yamin, Khurram
Wu, Steven
Wilder, Bryan
Machine Learning
Recent work has investigated the capabilities of large language models (LLMs) as zero-shot models for generating individual-level characteristics (e.g., to serve as risk models or augment survey datasets). However, when should a user have confidence that an LLM will provide high-quality predictions for their particular task? To address this question, we conduct a large-scale empirical study of LLMs' zero-shot predictive capabilities across a wide range of tabular prediction tasks. We find that LLMs' performance is highly variable, both on tasks within the same dataset and across different datasets. However, when the LLM performs well on the base prediction task, its predicted probabilities become a stronger signal for individual-level accuracy. Then, we construct metrics to predict LLMs' performance at the task level, aiming to distinguish between tasks where LLMs may perform well and where they are likely unsuitable. We find that some of these metrics, each of which are assessed without labeled data, yield strong signals of LLMs' predictive performance on new tasks.
title Predicting Language Models' Success at Zero-Shot Probabilistic Prediction
topic Machine Learning
url https://arxiv.org/abs/2509.15356