Hidden Topics: Measuring Sensitive AI Beliefs with List Experiments

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: Chupilkin, Maxim
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915816129167360
author Chupilkin, Maxim
author_facet Chupilkin, Maxim
contents How can researchers identify beliefs that large language models (LLMs) hide? As LLMs become more sophisticated and the prevalence of alignment faking increases, combined with their growing integration into high-stakes decision-making, responding to this challenge has become critical. This paper proposes that a list experiment, a simple method widely used in the social sciences, can be applied to study the hidden beliefs of LLMs. List experiments were originally developed to circumvent social desirability bias in human respondents, which closely parallels alignment faking in LLMs. The paper implements a list experiment on models developed by Anthropic, Google, and OpenAI and finds hidden approval of mass surveillance across all models, as well as some approval of torture, discrimination, and first nuclear strike. Importantly, a placebo treatment produces a null result, validating the method. The paper then compares list experiments with direct questioning and discusses the utility of the approach.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21939
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hidden Topics: Measuring Sensitive AI Beliefs with List Experiments
Chupilkin, Maxim
Computers and Society
Artificial Intelligence
How can researchers identify beliefs that large language models (LLMs) hide? As LLMs become more sophisticated and the prevalence of alignment faking increases, combined with their growing integration into high-stakes decision-making, responding to this challenge has become critical. This paper proposes that a list experiment, a simple method widely used in the social sciences, can be applied to study the hidden beliefs of LLMs. List experiments were originally developed to circumvent social desirability bias in human respondents, which closely parallels alignment faking in LLMs. The paper implements a list experiment on models developed by Anthropic, Google, and OpenAI and finds hidden approval of mass surveillance across all models, as well as some approval of torture, discrimination, and first nuclear strike. Importantly, a placebo treatment produces a null result, validating the method. The paper then compares list experiments with direct questioning and discusses the utility of the approach.
title Hidden Topics: Measuring Sensitive AI Beliefs with List Experiments
topic Computers and Society
Artificial Intelligence
url https://arxiv.org/abs/2602.21939