Measuring Scalar Constructs in Social Science with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Licht, Hauke, Sarkar, Rupak, Wu, Patrick Y., Goel, Pranav, Stoehr, Niklas, Ash, Elliott, Hoyle, Alexander Miserlis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918145407582208
author Licht, Hauke
Sarkar, Rupak
Wu, Patrick Y.
Goel, Pranav
Stoehr, Niklas
Ash, Elliott
Hoyle, Alexander Miserlis
author_facet Licht, Hauke
Sarkar, Rupak
Wu, Patrick Y.
Goel, Pranav
Stoehr, Niklas
Ash, Elliott
Hoyle, Alexander Miserlis
contents Many constructs that characterize language, like its complexity or emotionality, have a naturally continuous semantic structure; a public speech is not just "simple" or "complex," but exists on a continuum between extremes. Although large language models (LLMs) are an attractive tool for measuring scalar constructs, their idiosyncratic treatment of numerical outputs raises questions of how to best apply them. We address these questions with a comprehensive evaluation of LLM-based approaches to scalar construct measurement in social science. Using multiple datasets sourced from the political science literature, we evaluate four approaches: unweighted direct pointwise scoring, aggregation of pairwise comparisons, token-probability-weighted pointwise scoring, and finetuning. Our study finds that pairwise comparisons made by LLMs produce better measurements than simply prompting the LLM to directly output the scores, which suffers from bunching around arbitrary numbers. However, taking the weighted mean over the token probability of scores further improves the measurements over the two previous approaches. Finally, finetuning smaller models with as few as 1,000 training pairs can match or exceed the performance of prompted LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03116
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Measuring Scalar Constructs in Social Science with LLMs
Licht, Hauke
Sarkar, Rupak
Wu, Patrick Y.
Goel, Pranav
Stoehr, Niklas
Ash, Elliott
Hoyle, Alexander Miserlis
Computation and Language
Many constructs that characterize language, like its complexity or emotionality, have a naturally continuous semantic structure; a public speech is not just "simple" or "complex," but exists on a continuum between extremes. Although large language models (LLMs) are an attractive tool for measuring scalar constructs, their idiosyncratic treatment of numerical outputs raises questions of how to best apply them. We address these questions with a comprehensive evaluation of LLM-based approaches to scalar construct measurement in social science. Using multiple datasets sourced from the political science literature, we evaluate four approaches: unweighted direct pointwise scoring, aggregation of pairwise comparisons, token-probability-weighted pointwise scoring, and finetuning. Our study finds that pairwise comparisons made by LLMs produce better measurements than simply prompting the LLM to directly output the scores, which suffers from bunching around arbitrary numbers. However, taking the weighted mean over the token probability of scores further improves the measurements over the two previous approaches. Finally, finetuning smaller models with as few as 1,000 training pairs can match or exceed the performance of prompted LLMs.
title Measuring Scalar Constructs in Social Science with LLMs
topic Computation and Language
url https://arxiv.org/abs/2509.03116