Poor Alignment and Steerability of Large Language Models: Evidence from College Admission Essays

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Jinsook, Alvero, AJ, Joachims, Thorsten, Kizilcec, René
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912839171571712
author Lee, Jinsook
Alvero, AJ
Joachims, Thorsten
Kizilcec, René
author_facet Lee, Jinsook
Alvero, AJ
Joachims, Thorsten
Kizilcec, René
contents People are increasingly using technologies equipped with large language models (LLM) to write texts for formal communication, which raises two important questions at the intersection of technology and society: Who do LLMs write like (model alignment); and can LLMs be prompted to change who they write like (model steerability). We investigate these questions in the high-stakes context of undergraduate admissions at a selective university by comparing lexical and sentence variation between essays written by 30,000 applicants to two types of LLM-generated essays: one prompted with only the essay question used by the human applicants; and another with additional demographic information about each applicant. We consistently find that both types of LLM-generated essays are linguistically distinct from human-authored essays, regardless of the specific model and analytical approach. Further, prompting a specific sociodemographic identity is remarkably ineffective in aligning the model with the linguistic patterns observed in human writing from this identity group. This holds along the key dimensions of sex, race, first-generation status, and geographic location. The demographically prompted and unprompted synthetic texts were also more similar to each other than to the human text, meaning that prompting did not alleviate homogenization. These issues of model alignment and steerability in current LLMs raise concerns about the use of LLMs in high-stakes contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20062
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Poor Alignment and Steerability of Large Language Models: Evidence from College Admission Essays
Lee, Jinsook
Alvero, AJ
Joachims, Thorsten
Kizilcec, René
Computation and Language
People are increasingly using technologies equipped with large language models (LLM) to write texts for formal communication, which raises two important questions at the intersection of technology and society: Who do LLMs write like (model alignment); and can LLMs be prompted to change who they write like (model steerability). We investigate these questions in the high-stakes context of undergraduate admissions at a selective university by comparing lexical and sentence variation between essays written by 30,000 applicants to two types of LLM-generated essays: one prompted with only the essay question used by the human applicants; and another with additional demographic information about each applicant. We consistently find that both types of LLM-generated essays are linguistically distinct from human-authored essays, regardless of the specific model and analytical approach. Further, prompting a specific sociodemographic identity is remarkably ineffective in aligning the model with the linguistic patterns observed in human writing from this identity group. This holds along the key dimensions of sex, race, first-generation status, and geographic location. The demographically prompted and unprompted synthetic texts were also more similar to each other than to the human text, meaning that prompting did not alleviate homogenization. These issues of model alignment and steerability in current LLMs raise concerns about the use of LLMs in high-stakes contexts.
title Poor Alignment and Steerability of Large Language Models: Evidence from College Admission Essays
topic Computation and Language
url https://arxiv.org/abs/2503.20062