Human Psychometric Questionnaires Mischaracterize LLM Behavior
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Woojung, Choi, Dongmin, Park, Yoonah, Han, Jongwook, Lee, Eun-Ju, Jo, Yohan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
by: Han, Jongwook, et al.
Published: (2025)
by: Han, Jongwook, et al.
Published: (2025)
Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators
by: Lim, Sungjib, et al.
Published: (2025)
by: Lim, Sungjib, et al.
Published: (2025)
PVP: An Image Dataset for Personalized Visual Persuasion with Persuasion Strategies, Viewer Characteristics, and Persuasiveness Ratings
by: Kim, Junseo, et al.
Published: (2025)
by: Kim, Junseo, et al.
Published: (2025)
Quantifying Data Contamination in Psychometric Evaluations of LLMs
by: Han, Jongwook, et al.
Published: (2025)
by: Han, Jongwook, et al.
Published: (2025)
Improving Dialogue State Tracking through Combinatorial Search for In-Context Examples
by: Pyun, Haesung, et al.
Published: (2025)
by: Pyun, Haesung, et al.
Published: (2025)
Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction
by: Park, Sejun, et al.
Published: (2026)
by: Park, Sejun, et al.
Published: (2026)
Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
by: Lee, Jonggeun, et al.
Published: (2025)
by: Lee, Jonggeun, et al.
Published: (2025)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
by: Park, Junyeong, et al.
Published: (2025)
by: Park, Junyeong, et al.
Published: (2025)
Model-based Preference Optimization in Abstractive Summarization without Human Feedback
by: Choi, Jaepill, et al.
Published: (2024)
by: Choi, Jaepill, et al.
Published: (2024)
Context-Robust Knowledge Editing for Language Models
by: Park, Haewon, et al.
Published: (2025)
by: Park, Haewon, et al.
Published: (2025)
Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
by: Park, Yoonah, et al.
Published: (2025)
by: Park, Yoonah, et al.
Published: (2025)
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning
by: Lee, Gisang, et al.
Published: (2024)
by: Lee, Gisang, et al.
Published: (2024)
In-N-Out: A Parameter-Level API Graph Dataset for Tool Agents
by: Lee, Seungkyu, et al.
Published: (2025)
by: Lee, Seungkyu, et al.
Published: (2025)
Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
by: Kong, Injin, et al.
Published: (2026)
by: Kong, Injin, et al.
Published: (2026)
KL for a KL: On-Policy Distillation with Control Variate Baseline
by: Oh, Minjae, et al.
Published: (2026)
by: Oh, Minjae, et al.
Published: (2026)
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
by: Chae, Kyubyung, et al.
Published: (2024)
by: Chae, Kyubyung, et al.
Published: (2024)
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
by: Seo, Gyuhyeon, et al.
Published: (2025)
by: Seo, Gyuhyeon, et al.
Published: (2025)
Fakes of Varying Shades: How Warning Affects Human Perception and Engagement Regarding LLM Hallucinations
by: Nahar, Mahjabin, et al.
Published: (2024)
by: Nahar, Mahjabin, et al.
Published: (2024)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
by: Lee, Yooseop, et al.
Published: (2025)
by: Lee, Yooseop, et al.
Published: (2025)
Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models
by: Kong, Injin, et al.
Published: (2026)
by: Kong, Injin, et al.
Published: (2026)
Thinking Like a Doctor: Conversational Diagnosis through the Exploration of Diagnostic Knowledge Graphs
by: Won, Jeongmoon, et al.
Published: (2026)
by: Won, Jeongmoon, et al.
Published: (2026)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026)
by: Lee, Hyomin, et al.
Published: (2026)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models
by: Ye, Haoran, et al.
Published: (2024)
by: Ye, Haoran, et al.
Published: (2024)
Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning
by: Oh, Minjae, et al.
Published: (2025)
by: Oh, Minjae, et al.
Published: (2025)
Improving LLM Leaderboards with Psychometrical Methodology
by: Federiakin, Denis
Published: (2025)
by: Federiakin, Denis
Published: (2025)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
From Intuition to Calibrated Judgment: A Rubric-Based Expert-Panel Study of Human Detection of LLM-Generated Korean Text
by: Park, Shinwoo, et al.
Published: (2026)
by: Park, Shinwoo, et al.
Published: (2026)
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
by: Choi, Yunho, et al.
Published: (2026)
by: Choi, Yunho, et al.
Published: (2026)
KMI: A Dataset of Korean Motivational Interviewing Dialogues for Psychotherapy
by: Kim, Hyunjong, et al.
Published: (2025)
by: Kim, Hyunjong, et al.
Published: (2025)
Ever-Evolving Memory by Blending and Refining the Past
by: Kim, Seo Hyun, et al.
Published: (2024)
by: Kim, Seo Hyun, et al.
Published: (2024)
AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans
by: Xie, Wei, et al.
Published: (2025)
by: Xie, Wei, et al.
Published: (2025)
Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
by: Choi, Junhyuk, et al.
Published: (2025)
by: Choi, Junhyuk, et al.
Published: (2025)
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
by: Li, Peiyu, et al.
Published: (2025)
by: Li, Peiyu, et al.
Published: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
by: Kim, Heehwan, et al.
Published: (2025)
by: Kim, Heehwan, et al.
Published: (2025)
Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models
by: He-Yueya, Joy, et al.
Published: (2024)
by: He-Yueya, Joy, et al.
Published: (2024)
Scaling Bidirectional Spans and Span Violations in Attention Mechanism
by: Kim, Jongwook, et al.
Published: (2025)
by: Kim, Jongwook, et al.
Published: (2025)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
by: Moon, Sehwan, et al.
Published: (2025)
by: Moon, Sehwan, et al.
Published: (2025)
Database Normalization via Dual-LLM Self-Refinement
by: Jo, Eunjae, et al.
Published: (2025)
by: Jo, Eunjae, et al.
Published: (2025)
DialSim: A Dialogue Simulator for Evaluating Long-Term Multi-Party Dialogue Understanding of Conversational Agents
by: Kim, Jiho, et al.
Published: (2024)
by: Kim, Jiho, et al.
Published: (2024)
Similar Items
-
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
by: Han, Jongwook, et al.
Published: (2025) -
Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators
by: Lim, Sungjib, et al.
Published: (2025) -
PVP: An Image Dataset for Personalized Visual Persuasion with Persuasion Strategies, Viewer Characteristics, and Persuasiveness Ratings
by: Kim, Junseo, et al.
Published: (2025) -
Quantifying Data Contamination in Psychometric Evaluations of LLMs
by: Han, Jongwook, et al.
Published: (2025) -
Improving Dialogue State Tracking through Combinatorial Search for In-Context Examples
by: Pyun, Haesung, et al.
Published: (2025)