Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jingshen, Xie, Jiajun, Qiu, Xinying |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
von: Schmucker, Robin, et al.
Veröffentlicht: (2025)
von: Schmucker, Robin, et al.
Veröffentlicht: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
von: Wachter, Jasmin, et al.
Veröffentlicht: (2025)
von: Wachter, Jasmin, et al.
Veröffentlicht: (2025)
DualReward: A Dynamic Reinforcement Learning Framework for Cloze Tests Distractor Generation
von: Huang, Tianyou, et al.
Veröffentlicht: (2025)
von: Huang, Tianyou, et al.
Veröffentlicht: (2025)
Embedding Enhancement via Fine-Tuned Language Models for Learner-Item Cognitive Modeling
von: Liu, Yuanhao, et al.
Veröffentlicht: (2026)
von: Liu, Yuanhao, et al.
Veröffentlicht: (2026)
Difficulty-Controllable Cloze Question Distractor Generation
von: Kang, Seokhoon, et al.
Veröffentlicht: (2025)
von: Kang, Seokhoon, et al.
Veröffentlicht: (2025)
Constructing Cloze Questions Generatively
von: Sun, Yicheng, et al.
Veröffentlicht: (2024)
von: Sun, Yicheng, et al.
Veröffentlicht: (2024)
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
von: Liu, Yunting, et al.
Veröffentlicht: (2024)
von: Liu, Yunting, et al.
Veröffentlicht: (2024)
Automated Item Neutralization for Non-Cognitive Scales: A Large Language Model Approach to Reducing Social-Desirability Bias
von: Wu, Sirui, et al.
Veröffentlicht: (2025)
von: Wu, Sirui, et al.
Veröffentlicht: (2025)
Label Confidence Weighted Learning for Target-level Sentence Simplification
von: Qiu, Xinying, et al.
Veröffentlicht: (2024)
von: Qiu, Xinying, et al.
Veröffentlicht: (2024)
NLP Cluster Analysis of Common Core State Standards and NAEP Item Specifications
von: Camilli, Gregory, et al.
Veröffentlicht: (2024)
von: Camilli, Gregory, et al.
Veröffentlicht: (2024)
JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
von: Yao, Louie Hong, et al.
Veröffentlicht: (2025)
von: Yao, Louie Hong, et al.
Veröffentlicht: (2025)
Augmenting Rating-Scale Measures with Text-Derived Items Using the Information-Determined Scoring (IDS) Framework
von: Watson, Joe, et al.
Veröffentlicht: (2025)
von: Watson, Joe, et al.
Veröffentlicht: (2025)
Towards Enriched Controllability for Educational Question Generation
von: Leite, Bernardo, et al.
Veröffentlicht: (2023)
von: Leite, Bernardo, et al.
Veröffentlicht: (2023)
Length-Induced Embedding Collapse in PLM-based Models
von: Zhou, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2024)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
von: Wang, Tianyu, et al.
Veröffentlicht: (2026)
von: Wang, Tianyu, et al.
Veröffentlicht: (2026)
QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation
von: Nguyen, Bang, et al.
Veröffentlicht: (2025)
von: Nguyen, Bang, et al.
Veröffentlicht: (2025)
Reporting and Analysing the Environmental Impact of Language Models on the Example of Commonsense Question Answering with External Knowledge
von: Usmanova, Aida, et al.
Veröffentlicht: (2024)
von: Usmanova, Aida, et al.
Veröffentlicht: (2024)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2025)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2025)
On the Credibility of Evaluating LLMs using Survey Questions
von: Libovický, Jindřich
Veröffentlicht: (2026)
von: Libovický, Jindřich
Veröffentlicht: (2026)
Social Bias in Popular Question-Answering Benchmarks
von: Kraft, Angelie, et al.
Veröffentlicht: (2025)
von: Kraft, Angelie, et al.
Veröffentlicht: (2025)
Understanding Social Support Needs in Questions: A Hybrid Approach Integrating Semi-Supervised Learning and LLM-based Data Augmentation
von: Kuang, Junwei, et al.
Veröffentlicht: (2025)
von: Kuang, Junwei, et al.
Veröffentlicht: (2025)
EduAgentQG: A Multi-Agent Workflow Framework for Personalized Question Generation
von: Jia, Rui, et al.
Veröffentlicht: (2025)
von: Jia, Rui, et al.
Veröffentlicht: (2025)
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
von: Atif, Farah, et al.
Veröffentlicht: (2025)
von: Atif, Farah, et al.
Veröffentlicht: (2025)
Psychological Assessments with Large Language Models: A Privacy-Focused and Cost-Effective Approach
von: Blanco-Cuaresma, Sergi
Veröffentlicht: (2024)
von: Blanco-Cuaresma, Sergi
Veröffentlicht: (2024)
Evaluating Psychological Safety of Large Language Models
von: Li, Xingxuan, et al.
Veröffentlicht: (2022)
von: Li, Xingxuan, et al.
Veröffentlicht: (2022)
Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
von: Sun, Yuxi, et al.
Veröffentlicht: (2025)
von: Sun, Yuxi, et al.
Veröffentlicht: (2025)
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral
von: Steenhuis, Quinten, et al.
Veröffentlicht: (2026)
von: Steenhuis, Quinten, et al.
Veröffentlicht: (2026)
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
von: Singh, Shrutika, et al.
Veröffentlicht: (2025)
von: Singh, Shrutika, et al.
Veröffentlicht: (2025)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2024)
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2024)
Large Language Models Leverage External Knowledge to Extend Clinical Insight Beyond Language Boundaries
von: Wu, Jiageng, et al.
Veröffentlicht: (2023)
von: Wu, Jiageng, et al.
Veröffentlicht: (2023)
Synthetic Student Responses: LLM-Extracted Features for IRT Difficulty Parameter Estimation
von: Hoyl, Matias
Veröffentlicht: (2026)
von: Hoyl, Matias
Veröffentlicht: (2026)
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
von: Duan, Shitong, et al.
Veröffentlicht: (2023)
von: Duan, Shitong, et al.
Veröffentlicht: (2023)
When AI Meets Early Childhood Education: Large Language Models as Assessment Teammates in Chinese Preschools
von: Li, Xingming, et al.
Veröffentlicht: (2026)
von: Li, Xingming, et al.
Veröffentlicht: (2026)
Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
von: Haider, Batool, et al.
Veröffentlicht: (2025)
von: Haider, Batool, et al.
Veröffentlicht: (2025)
Implicit assessment of language learning during practice as accurate as explicit testing
von: Hou, Jue, et al.
Veröffentlicht: (2024)
von: Hou, Jue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
von: Schmucker, Robin, et al.
Veröffentlicht: (2025) -
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
von: Li, Ming, et al.
Veröffentlicht: (2025) -
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
von: Wachter, Jasmin, et al.
Veröffentlicht: (2025) -
DualReward: A Dynamic Reinforcement Learning Framework for Cloze Tests Distractor Generation
von: Huang, Tianyou, et al.
Veröffentlicht: (2025) -
Embedding Enhancement via Fine-Tuned Language Models for Learner-Item Cognitive Modeling
von: Liu, Yuanhao, et al.
Veröffentlicht: (2026)