"Mirror" Language AI Models of Depression are Criterion-Contaminated

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Tong, Hussain, Rasiq, Gupta, Mehak, Oltmanns, Joshua R.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911218096144384
author Li, Tong
Hussain, Rasiq
Gupta, Mehak
Oltmanns, Joshua R.
author_facet Li, Tong
Hussain, Rasiq
Gupta, Mehak
Oltmanns, Joshua R.
contents Recent studies show near-perfect language-based predictions of depression scores (R2 = .70), but these "Mirror" models rely on language responses directly from depression assessments to predict depression assessment scores. These methods suffer from criterion contamination that inflate prediction estimates. We compare "Mirror" models to "Non-Mirror" models, which use other external language to predict depression scores. 110 participants completed both structured diagnostic (Mirror condition) and life history (Non-Mirror condition) interviews. LLMs were prompted to predict diagnostic depression scores. As expected, Mirror models were near-perfect. However, Non-Mirror models also displayed prediction sizes considered large in psychology. Further, both Mirror and Non-Mirror predictions correlated with other questionnaire-based depression symptoms at similar sizes, suggesting bias in Mirror models. Topic modeling revealed different theme structures across model types. As language models for depression continue to evolve, incorporating Non-Mirror approaches may support more valid and clinically useful language-based AI applications in psychological assessment.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05830
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle "Mirror" Language AI Models of Depression are Criterion-Contaminated
Li, Tong
Hussain, Rasiq
Gupta, Mehak
Oltmanns, Joshua R.
Computation and Language
Computers and Society
Recent studies show near-perfect language-based predictions of depression scores (R2 = .70), but these "Mirror" models rely on language responses directly from depression assessments to predict depression assessment scores. These methods suffer from criterion contamination that inflate prediction estimates. We compare "Mirror" models to "Non-Mirror" models, which use other external language to predict depression scores. 110 participants completed both structured diagnostic (Mirror condition) and life history (Non-Mirror condition) interviews. LLMs were prompted to predict diagnostic depression scores. As expected, Mirror models were near-perfect. However, Non-Mirror models also displayed prediction sizes considered large in psychology. Further, both Mirror and Non-Mirror predictions correlated with other questionnaire-based depression symptoms at similar sizes, suggesting bias in Mirror models. Topic modeling revealed different theme structures across model types. As language models for depression continue to evolve, incorporating Non-Mirror approaches may support more valid and clinically useful language-based AI applications in psychological assessment.
title "Mirror" Language AI Models of Depression are Criterion-Contaminated
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2508.05830