Do LLMs Give Psychometrically Plausible Responses in Educational Assessments?
Fuente:
arXiv
Saved in:
| Main Authors: | Säuberli, Andreas, Frassinelli, Diego, Plank, Barbara |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Controlling Reading Ease with Gaze-Guided Text Generation
by: Säuberli, Andreas, et al.
Published: (2026)
by: Säuberli, Andreas, et al.
Published: (2026)
Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
by: Lan, Jian, et al.
Published: (2024)
by: Lan, Jian, et al.
Published: (2024)
Resource-Lean Lexicon Induction for German Dialects
by: Litschko, Robert, et al.
Published: (2026)
by: Litschko, Robert, et al.
Published: (2026)
I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes
by: Zhou, Shijia, et al.
Published: (2026)
by: Zhou, Shijia, et al.
Published: (2026)
Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora
by: Litschko, Robert, et al.
Published: (2025)
by: Litschko, Robert, et al.
Published: (2025)
To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity
by: Sedova, Anastasiia, et al.
Published: (2024)
by: Sedova, Anastasiia, et al.
Published: (2024)
Automatic Generation and Evaluation of Reading Comprehension Test Items with Large Language Models
by: Säuberli, Andreas, et al.
Published: (2024)
by: Säuberli, Andreas, et al.
Published: (2024)
Generalizable Sarcasm Detection Is Just Around The Corner, Of Course!
by: Jang, Hyewon, et al.
Published: (2024)
by: Jang, Hyewon, et al.
Published: (2024)
Digital Comprehensibility Assessment of Simplified Texts among Persons with Intellectual Disabilities
by: Säuberli, Andreas, et al.
Published: (2024)
by: Säuberli, Andreas, et al.
Published: (2024)
Investigating the Nature of Disagreements on Mid-Scale Ratings: A Case Study on the Abstractness-Concreteness Continuum
by: Knupleš, Urban, et al.
Published: (2023)
by: Knupleš, Urban, et al.
Published: (2023)
RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
by: Testoni, Alberto, et al.
Published: (2024)
by: Testoni, Alberto, et al.
Published: (2024)
LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task
by: Leonardelli, Elisa, et al.
Published: (2025)
by: Leonardelli, Elisa, et al.
Published: (2025)
Unveiling the Mystery of Visual Attributes of Concrete and Abstract Concepts: Variability, Nearest Neighbors, and Challenging Categories
by: Tater, Tarun, et al.
Published: (2024)
by: Tater, Tarun, et al.
Published: (2024)
Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models
by: Ahnert, Georg, et al.
Published: (2025)
by: Ahnert, Georg, et al.
Published: (2025)
What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility
by: Wang, Sheng-Fu, et al.
Published: (2025)
by: Wang, Sheng-Fu, et al.
Published: (2025)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
by: Lee, Seungbeen, et al.
Published: (2024)
by: Lee, Seungbeen, et al.
Published: (2024)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
B4: Towards Optimal Assessment of Plausible Code Solutions with Plausible Tests
by: Chen, Mouxiang, et al.
Published: (2024)
by: Chen, Mouxiang, et al.
Published: (2024)
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
by: Eichin, Florian, et al.
Published: (2025)
by: Eichin, Florian, et al.
Published: (2025)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
PRobELM: Plausibility Ranking Evaluation for Language Models
by: Yuan, Zhangdie, et al.
Published: (2024)
by: Yuan, Zhangdie, et al.
Published: (2024)
Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not
by: Karakaş, Sercan
Published: (2026)
by: Karakaş, Sercan
Published: (2026)
MaiBaam Annotation Guidelines
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
by: Orth, Jasmin, et al.
Published: (2025)
by: Orth, Jasmin, et al.
Published: (2025)
EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
by: Zuo, Longfei, et al.
Published: (2025)
by: Zuo, Longfei, et al.
Published: (2025)
Add Noise, Tasks, or Layers? MaiNLP at the VarDial 2025 Shared Task on Norwegian Dialectal Slot and Intent Detection
by: Blaschke, Verena, et al.
Published: (2025)
by: Blaschke, Verena, et al.
Published: (2025)
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
by: Blaschke, Verena, et al.
Published: (2025)
by: Blaschke, Verena, et al.
Published: (2025)
Indirect Question Answering in English, German and Bavarian: A Challenging Task for High- and Low-Resource Languages Alike
by: Winkler, Miriam, et al.
Published: (2026)
by: Winkler, Miriam, et al.
Published: (2026)
CLIMATELI: Evaluating Entity Linking on Climate Change Data
by: Zhou, Shijia, et al.
Published: (2024)
by: Zhou, Shijia, et al.
Published: (2024)
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German Varieties
by: Artemova, Ekaterina, et al.
Published: (2024)
by: Artemova, Ekaterina, et al.
Published: (2024)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
by: Si, Shengyun, et al.
Published: (2025)
by: Si, Shengyun, et al.
Published: (2025)
Better Aligned with Survey Respondents or Training Data? Unveiling Political Leanings of LLMs on U.S. Supreme Court Cases
by: Xu, Shanshan, et al.
Published: (2025)
by: Xu, Shanshan, et al.
Published: (2025)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
by: Chen, Beiduo, et al.
Published: (2024)
by: Chen, Beiduo, et al.
Published: (2024)
Plausibility Vaccine: Injecting LLM Knowledge for Event Plausibility
by: Chmura, Jacob, et al.
Published: (2025)
by: Chmura, Jacob, et al.
Published: (2025)
AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility Estimation
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal Bavarian Case Study
by: Krückl, Xaver Maria, et al.
Published: (2025)
by: Krückl, Xaver Maria, et al.
Published: (2025)
Similar Items
-
Controlling Reading Ease with Gaze-Guided Text Generation
by: Säuberli, Andreas, et al.
Published: (2026) -
Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
by: Lan, Jian, et al.
Published: (2024) -
Resource-Lean Lexicon Induction for German Dialects
by: Litschko, Robert, et al.
Published: (2026) -
I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes
by: Zhou, Shijia, et al.
Published: (2026) -
Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora
by: Litschko, Robert, et al.
Published: (2025)