You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments
Fuente:
arXiv
Saved in:
| Main Authors: | Shu, Bangzhao, Zhang, Lechen, Choi, Minje, Dunagan, Lavinia, Logeswaran, Lajanugen, Lee, Moontae, Card, Dallas, Jurgens, David |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
by: Zhang, Lechen, et al.
Published: (2024)
by: Zhang, Lechen, et al.
Published: (2024)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
by: Zhang, Lechen, et al.
Published: (2025)
by: Zhang, Lechen, et al.
Published: (2025)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
by: Zheng, Mingqian, et al.
Published: (2023)
by: Zheng, Mingqian, et al.
Published: (2023)
“You don't know what you don't know”: A qualitative study of informational needs of patients, family members, and living donors to inform transplant system metrics
by: Allyson Hart, et al.
Published: (2024)
by: Allyson Hart, et al.
Published: (2024)
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus
by: Litterer, Benjamin, et al.
Published: (2024)
by: Litterer, Benjamin, et al.
Published: (2024)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
by: Khalifa, Muhammad, et al.
Published: (2023)
by: Khalifa, Muhammad, et al.
Published: (2023)
Multicancer detection tests: What we know and what we don’t know
by: Sam M. Hanash, et al.
Published: (2024)
by: Sam M. Hanash, et al.
Published: (2024)
‘I don't know’—reclaiming not‐knowing in medical transitions
by: Yvonne Carlsson, et al.
Published: (2025)
by: Yvonne Carlsson, et al.
Published: (2025)
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
by: Shen, Siqi, et al.
Published: (2025)
by: Shen, Siqi, et al.
Published: (2025)
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
by: Shen, Siqi, et al.
Published: (2024)
by: Shen, Siqi, et al.
Published: (2024)
“I know what I don't know”: Metacognition in leadership learning
by: Jillian Volpe‐White
Published: (2024)
by: Jillian Volpe‐White
Published: (2024)
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024)
by: Aitchison, Laurence
Published: (2024)
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
by: Zhang, Yunxiang, et al.
Published: (2024)
by: Zhang, Yunxiang, et al.
Published: (2024)
Wolf-Rayet stars -- what we know and what we don't
by: Maryeva, Olga
Published: (2024)
by: Maryeva, Olga
Published: (2024)
Xeno Amino Acids: A look into biochemistry as we don't know it
by: Brown, Sean M., et al.
Published: (2023)
by: Brown, Sean M., et al.
Published: (2023)
Fetal Leydig cells: What we know and what we don't
by: Keer Jiang, et al.
Published: (2024)
by: Keer Jiang, et al.
Published: (2024)
Gut size flexibility in rodents: what we know, and don’t know, after a century of research
by: DANIEL E. NAYA
Published: (2008)
by: DANIEL E. NAYA
Published: (2008)
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)
by: Khalifa, Muhammad, et al.
Published: (2025)
Why we still don't know much about housing supply elasticity
by: Daniel A. Broxterman, et al.
Published: (2026)
by: Daniel A. Broxterman, et al.
Published: (2026)
SoK: What don't we know? Understanding Security Vulnerabilities in SNARKs
by: Chaliasos, Stefanos, et al.
Published: (2024)
by: Chaliasos, Stefanos, et al.
Published: (2024)
“It's like an epidemic, we don't know what to do”: The perceived need for and benefits of a suicide prevention programme in UK schools
by: Emma Ashworth, et al.
Published: (2024)
by: Emma Ashworth, et al.
Published: (2024)
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Economic inequality and social mobility in preindustrial societies: What we know, what we don't (but should) know
by: Guido Alfani
Published: (2026)
by: Guido Alfani
Published: (2026)
Analyzing the Engagement of Social Relationships During Life Event Shocks in Social Media
by: Choi, Minje, et al.
Published: (2023)
by: Choi, Minje, et al.
Published: (2023)
What we still don't know about weed diversity: A scoping review
by: Anna S. Westbrook, et al.
Published: (2024)
by: Anna S. Westbrook, et al.
Published: (2024)
On the consistent reasoning paradox of intelligence and optimal trust in AI: The power of 'I don't know'
by: Bastounis, Alexander, et al.
Published: (2024)
by: Bastounis, Alexander, et al.
Published: (2024)
Green growth, just transition, and green jobs: there's a lot we don't know
by: International Labour Organization. Employment Policy Department., et al.
Published: (2018)
by: International Labour Organization. Employment Policy Department., et al.
Published: (2018)
Neglected seed dispersers and research compartmentalisation: how much do we know about what we don't know?
by: Sara Beatriz Mendes, et al.
Published: (2026)
by: Sara Beatriz Mendes, et al.
Published: (2026)
Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark
by: Choi, Minje, et al.
Published: (2023)
by: Choi, Minje, et al.
Published: (2023)
Boys don’t cry!
Published: (2019)
Published: (2019)
Show, don't tell
Published: (2021)
Published: (2021)
"If we don't they won't"
by: Ehrhardt, Harryette B.
Published: (1969)
by: Ehrhardt, Harryette B.
Published: (1969)
Hyperhidrosis: don't sweat it
by: Mitchell J. Lycett, et al.
Published: (2025)
by: Mitchell J. Lycett, et al.
Published: (2025)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
by: Khalifa, Muhammad, et al.
Published: (2026)
by: Khalifa, Muhammad, et al.
Published: (2026)
Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs
by: Sun, Huaman, et al.
Published: (2023)
by: Sun, Huaman, et al.
Published: (2023)
Collective Memory and Narrative Cohesion: A Computational Study of Palestinian Refugee Oral Histories in Lebanon
by: Awwad, Ghadeer, et al.
Published: (2025)
by: Awwad, Ghadeer, et al.
Published: (2025)
"You don't need a university degree to comprehend data protection this way": LLM-Powered Interactive Privacy Policy Assessment
by: Freiberger, Vincent, et al.
Published: (2025)
by: Freiberger, Vincent, et al.
Published: (2025)
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
For those who don't know (how) to ask: Building a dataset of technology questions for digital newcomers
by: Lucas, Evan, et al.
Published: (2024)
by: Lucas, Evan, et al.
Published: (2024)
Shadows Don't Lie and Lines Can't Bend! Generative Models don't know Projective Geometry...for now
by: Sarkar, Ayush, et al.
Published: (2023)
by: Sarkar, Ayush, et al.
Published: (2023)
Similar Items
-
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
by: Zhang, Lechen, et al.
Published: (2024) -
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
by: Zhang, Lechen, et al.
Published: (2025) -
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
by: Zheng, Mingqian, et al.
Published: (2023) -
“You don't know what you don't know”: A qualitative study of informational needs of patients, family members, and living donors to inform transplant system metrics
by: Allyson Hart, et al.
Published: (2024) -
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus
by: Litterer, Benjamin, et al.
Published: (2024)