Saved in:
| Main Authors: | Kapoor, Radhika, Truong, Sang T., Haber, Nick, Ruiz-Primo, Maria Araceli, Domingue, Benjamin W. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.20663 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025)
by: Schmucker, Robin, et al.
Published: (2025)
A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation
by: Hwang, Seonjeong, et al.
Published: (2026)
by: Hwang, Seonjeong, et al.
Published: (2026)
Using Vision + Language Models to Predict Item Difficulty
by: Khan, Samin
Published: (2026)
by: Khan, Samin
Published: (2026)
Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
by: Hwang, Seonjeong, et al.
Published: (2025)
by: Hwang, Seonjeong, et al.
Published: (2025)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
Automatic Generation and Evaluation of Reading Comprehension Test Items with Large Language Models
by: Säuberli, Andreas, et al.
Published: (2024)
by: Säuberli, Andreas, et al.
Published: (2024)
RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
by: Li, Xinyuan, et al.
Published: (2025)
by: Li, Xinyuan, et al.
Published: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Items contextualizados de ciencias en PISA: la conexión entre las demandas cognitivas y las características de contexto de los items
by: Maria-Araceli Ruiz-Primo
Published: (2016)
by: Maria-Araceli Ruiz-Primo
Published: (2016)
Can Model Uncertainty Function as a Proxy for Multiple-Choice Question Item Difficulty?
by: Zotos, Leonidas, et al.
Published: (2024)
by: Zotos, Leonidas, et al.
Published: (2024)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
by: Acquaye, Christabel, et al.
Published: (2026)
by: Acquaye, Christabel, et al.
Published: (2026)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
Polytomous Explanatory Item Response Models for Item Discrimination: Assessing Negative-Framing Effects in Social-Emotional Learning Surveys
by: Gilbert, Joshua B., et al.
Published: (2024)
by: Gilbert, Joshua B., et al.
Published: (2024)
Estimating Heterogeneous Treatment Effects with Item-Level Outcome Data: Insights from Item Response Theory
by: Gilbert, Joshua B., et al.
Published: (2024)
by: Gilbert, Joshua B., et al.
Published: (2024)
Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms
by: Razavi, Pooya, et al.
Published: (2025)
by: Razavi, Pooya, et al.
Published: (2025)
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
by: Zhang, Jingshen, et al.
Published: (2024)
by: Zhang, Jingshen, et al.
Published: (2024)
Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests
by: Fu, Yanbin, et al.
Published: (2025)
by: Fu, Yanbin, et al.
Published: (2025)
Item-Language Model for Conversational Recommendation
by: Yang, Li, et al.
Published: (2024)
by: Yang, Li, et al.
Published: (2024)
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Auditing LLM Benchmarks with Item Response Theory
by: Land, Sander, et al.
Published: (2026)
by: Land, Sander, et al.
Published: (2026)
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
by: Cong, Longwei, et al.
Published: (2026)
by: Cong, Longwei, et al.
Published: (2026)
Dynamic Bayesian Item Response Model with Decomposition (D-BIRD): Modeling Cohort and Individual Learning Over Time
by: Lee, Hansol, et al.
Published: (2025)
by: Lee, Hansol, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation
by: Lin, Fan, et al.
Published: (2024)
by: Lin, Fan, et al.
Published: (2024)
RAGSys: Item-Cold-Start Recommender as RAG System
by: Contal, Emile, et al.
Published: (2024)
by: Contal, Emile, et al.
Published: (2024)
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
by: Liu, Yunting, et al.
Published: (2024)
by: Liu, Yunting, et al.
Published: (2024)
From Past To Path: Masked History Learning for Next-Item Prediction in Generative Recommendation
by: Wei, KaiWen, et al.
Published: (2025)
by: Wei, KaiWen, et al.
Published: (2025)
Enriching Social Science Research via Survey Item Linking
by: Tsereteli, Tornike, et al.
Published: (2024)
by: Tsereteli, Tornike, et al.
Published: (2024)
Semantic Cells: Evolutional Process to Acquire Sense Diversity of Items
by: Ohsawa, Yukio, et al.
Published: (2024)
by: Ohsawa, Yukio, et al.
Published: (2024)
Action-Item-Driven Summarization of Long Meeting Transcripts
by: Golia, Logan, et al.
Published: (2023)
by: Golia, Logan, et al.
Published: (2023)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
by: Patel, Fagun, et al.
Published: (2025)
by: Patel, Fagun, et al.
Published: (2025)
Break Out the Silverware -- Semantic Understanding of Stored Household Items
by: Levi-Richter, Michaela, et al.
Published: (2025)
by: Levi-Richter, Michaela, et al.
Published: (2025)
Difficulty as a Proxy for Measuring Intrinsic Cognitive Load Item
by: Cai, Minghao, et al.
Published: (2025)
by: Cai, Minghao, et al.
Published: (2025)
Finding Words Associated with DIF: Predicting Differential Item Functioning using LLMs and Explainable AI
by: Maeda, Hotaka, et al.
Published: (2025)
by: Maeda, Hotaka, et al.
Published: (2025)
Multi-Agent Collaborative Filtering: Orchestrating Users and Items for Agentic Recommendations
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
by: Luo, Zhimeng, et al.
Published: (2025)
by: Luo, Zhimeng, et al.
Published: (2025)
Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators
by: Lim, Sungjib, et al.
Published: (2025)
by: Lim, Sungjib, et al.
Published: (2025)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
by: Zhou, Hongli, et al.
Published: (2025)
by: Zhou, Hongli, et al.
Published: (2025)
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs
by: Feucht, Sheridan, et al.
Published: (2024)
by: Feucht, Sheridan, et al.
Published: (2024)
Similar Items
-
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025) -
A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation
by: Hwang, Seonjeong, et al.
Published: (2026) -
Using Vision + Language Models to Predict Item Difficulty
by: Khan, Samin
Published: (2026) -
Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
by: Hwang, Seonjeong, et al.
Published: (2025) -
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025)