Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ming, Chen, Han, Xiao, Yunze, Chen, Jian, Jiao, Hong, Zhou, Tianyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025)
by: Schmucker, Robin, et al.
Published: (2025)
Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms
by: Razavi, Pooya, et al.
Published: (2025)
by: Razavi, Pooya, et al.
Published: (2025)
Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus
by: Brant, Thiago, et al.
Published: (2026)
by: Brant, Thiago, et al.
Published: (2026)
Detecting Struggling Student Programmers using Proficiency Taxonomies
by: Schwartz, Noga, et al.
Published: (2025)
by: Schwartz, Noga, et al.
Published: (2025)
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
by: Zhang, Jingshen, et al.
Published: (2024)
by: Zhang, Jingshen, et al.
Published: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Does Difficulty even Matter? Investigating Difficulty Adjustment and Practice Behavior in an Open-Ended Learning Task
by: Schütt, Anan, et al.
Published: (2023)
by: Schütt, Anan, et al.
Published: (2023)
Active Learning to Guide Labeling Efforts for Question Difficulty Estimation
by: Thuy, Arthur, et al.
Published: (2024)
by: Thuy, Arthur, et al.
Published: (2024)
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals
by: Ehara, Yo
Published: (2026)
by: Ehara, Yo
Published: (2026)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
by: Acquaye, Christabel, et al.
Published: (2026)
by: Acquaye, Christabel, et al.
Published: (2026)
The Course Difficulty Analysis Cookbook
by: Baucks, Frederik, et al.
Published: (2025)
by: Baucks, Frederik, et al.
Published: (2025)
PS$^2$: Parameterized Control for Fine-Grained Student Proficiency Simulation
by: Liu, Ruochen, et al.
Published: (2026)
by: Liu, Ruochen, et al.
Published: (2026)
Synthetic Student Responses: LLM-Extracted Features for IRT Difficulty Parameter Estimation
by: Hoyl, Matias
Published: (2026)
by: Hoyl, Matias
Published: (2026)
Prediction of Item Difficulty for Reading Comprehension Items by Creation of Annotated Item Repository
by: Kapoor, Radhika, et al.
Published: (2025)
by: Kapoor, Radhika, et al.
Published: (2025)
Determining the Difficulties of Students With Dyslexia via Virtual Reality and Artificial Intelligence: An Exploratory Analysis
by: Yeguas-Bolívar, Enrique, et al.
Published: (2024)
by: Yeguas-Bolívar, Enrique, et al.
Published: (2024)
Embracing Contradiction: Theoretical Inconsistency Will Not Impede the Road of Building Responsible AI Systems
by: Dai, Gordon, et al.
Published: (2025)
by: Dai, Gordon, et al.
Published: (2025)
Effects of Generative AI Errors on User Reliance Across Task Difficulty
by: Anthis, Jacy Reese, et al.
Published: (2026)
by: Anthis, Jacy Reese, et al.
Published: (2026)
The Use of Computational Thinking Skills, Difficulties, and Strategies of Introductory Programming Students Solving Bebras Tasks
by: Benedetti, Enrico, et al.
Published: (2026)
by: Benedetti, Enrico, et al.
Published: (2026)
Using Vision + Language Models to Predict Item Difficulty
by: Khan, Samin
Published: (2026)
by: Khan, Samin
Published: (2026)
From ChatGPT to DeepSeek: Can LLMs Simulate Humanity?
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Which Type of Students can LLMs Act? Investigating Authentic Simulation with Graph-based Human-AI Collaborative System
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
by: Xiao, Yang, et al.
Published: (2023)
by: Xiao, Yang, et al.
Published: (2023)
ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle
by: Miroyan, Mihran, et al.
Published: (2025)
by: Miroyan, Mihran, et al.
Published: (2025)
When AI Navigates the Fog of War
by: Li, Ming, et al.
Published: (2026)
by: Li, Ming, et al.
Published: (2026)
Validated Hypotheses as a Lens for Human-Likeness Evaluation in AI Agents
by: Liu, Xuan, et al.
Published: (2026)
by: Liu, Xuan, et al.
Published: (2026)
MCQ Difficulty Prediction via Modeling Learner Heterogeneity Using Data-Driven Cognitive Profiling
by: Krishnan, Dhriti, et al.
Published: (2026)
by: Krishnan, Dhriti, et al.
Published: (2026)
Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2026)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2026)
Can Model Uncertainty Function as a Proxy for Multiple-Choice Question Item Difficulty?
by: Zotos, Leonidas, et al.
Published: (2024)
by: Zotos, Leonidas, et al.
Published: (2024)
Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
by: Chakrabarty, Tuhin, et al.
Published: (2024)
by: Chakrabarty, Tuhin, et al.
Published: (2024)
Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
by: Li, Ming, et al.
Published: (2026)
by: Li, Ming, et al.
Published: (2026)
LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents
by: Xiao, Chang, et al.
Published: (2024)
by: Xiao, Chang, et al.
Published: (2024)
Gaining Insights into Group-Level Course Difficulty via Differential Course Functioning
by: Baucks, Frederik, et al.
Published: (2024)
by: Baucks, Frederik, et al.
Published: (2024)
Representative Social Choice: From Learning Theory to AI Alignment
by: Qiu, Tianyi
Published: (2024)
by: Qiu, Tianyi
Published: (2024)
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
by: Alipour, Shayan, et al.
Published: (2024)
by: Alipour, Shayan, et al.
Published: (2024)
Struggle Premium : How Human Effort and Imperfection Drive Perceived Value in the Age of AI
by: Sultana, Nazneen, et al.
Published: (2026)
by: Sultana, Nazneen, et al.
Published: (2026)
QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation
by: Nguyen, Bang, et al.
Published: (2025)
by: Nguyen, Bang, et al.
Published: (2025)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
Assessing Simulation Knowledge and Proficiency Among Undergraduate Computing Students in Brazil: Insights and Results from a Survey Research
by: Rodrigues, Fernando Brito, et al.
Published: (2025)
by: Rodrigues, Fernando Brito, et al.
Published: (2025)
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
by: Chen, Benjamin Minhao, et al.
Published: (2026)
by: Chen, Benjamin Minhao, et al.
Published: (2026)
Similar Items
-
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025) -
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025) -
Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms
by: Razavi, Pooya, et al.
Published: (2025) -
Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus
by: Brant, Thiago, et al.
Published: (2026) -
Detecting Struggling Student Programmers using Proficiency Taxonomies
by: Schwartz, Noga, et al.
Published: (2025)