Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
Fuente:
arXiv
Saved in:
| Main Authors: | Acquaye, Christabel, Huang, Yi Ting, Carpuat, Marine, Rudinger, Rachel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the U.S
by: Acquaye, Christabel, et al.
Published: (2024)
by: Acquaye, Christabel, et al.
Published: (2024)
Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender?
by: An, Haozhe, et al.
Published: (2024)
by: An, Haozhe, et al.
Published: (2024)
How often are errors in natural language reasoning due to paraphrastic variability?
by: Srikanth, Neha, et al.
Published: (2024)
by: Srikanth, Neha, et al.
Published: (2024)
Multiple LLM Agents Debate for Equitable Cultural Alignment
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
by: Palta, Shramay, et al.
Published: (2024)
by: Palta, Shramay, et al.
Published: (2024)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
AskQE: Question Answering as Automatic Evaluation for Machine Translation
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
My LLM might Mimic AAE -- But When Should it?
by: Sandoval, Sandra C., et al.
Published: (2025)
by: Sandoval, Sandra C., et al.
Published: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Steering Large Language Models with Register Analysis for Arbitrary Style Transfer
by: Yang, Xinchen, et al.
Published: (2025)
by: Yang, Xinchen, et al.
Published: (2025)
Do Text Simplification Systems Preserve Meaning? A Human Evaluation via Reading Comprehension
by: Agrawal, Sweta, et al.
Published: (2023)
by: Agrawal, Sweta, et al.
Published: (2023)
SpeechQE: Estimating the Quality of Direct Speech Translation
by: Han, HyoJung, et al.
Published: (2024)
by: Han, HyoJung, et al.
Published: (2024)
Should We be Pedantic About Reasoning Errors in Machine Translation?
by: Bao, Calvin, et al.
Published: (2026)
by: Bao, Calvin, et al.
Published: (2026)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
by: Ki, Dayeon, et al.
Published: (2024)
by: Ki, Dayeon, et al.
Published: (2024)
How Multilingual Are Large Language Models Fine-Tuned for Translation?
by: Richburg, Aquia, et al.
Published: (2024)
by: Richburg, Aquia, et al.
Published: (2024)
Automatic Input Rewriting Improves Translation with Large Language Models
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Keep It Private: Unsupervised Privatization of Online Text
by: Bao, Calvin, et al.
Published: (2024)
by: Bao, Calvin, et al.
Published: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
by: Srikanth, Neha, et al.
Published: (2026)
by: Srikanth, Neha, et al.
Published: (2026)
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
by: Hou, Yu, et al.
Published: (2025)
by: Hou, Yu, et al.
Published: (2025)
On the Mutual Influence of Gender and Occupation in LLM Representations
by: An, Haozhe, et al.
Published: (2025)
by: An, Haozhe, et al.
Published: (2025)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
by: Srikanth, Neha, et al.
Published: (2025)
by: Srikanth, Neha, et al.
Published: (2025)
Words as Bridges: Exploring Computational Support for Cross-Disciplinary Translation Work
by: Bao, Calvin, et al.
Published: (2025)
by: Bao, Calvin, et al.
Published: (2025)
Can Model Uncertainty Function as a Proxy for Multiple-Choice Question Item Difficulty?
by: Zotos, Leonidas, et al.
Published: (2024)
by: Zotos, Leonidas, et al.
Published: (2024)
Prediction of Item Difficulty for Reading Comprehension Items by Creation of Annotated Item Repository
by: Kapoor, Radhika, et al.
Published: (2025)
by: Kapoor, Radhika, et al.
Published: (2025)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
by: Cong, Longwei, et al.
Published: (2026)
by: Cong, Longwei, et al.
Published: (2026)
Can you map it to English? The Role of Cross-Lingual Alignment in Multilingual Performance of LLMs
by: Ravisankar, Kartik, et al.
Published: (2025)
by: Ravisankar, Kartik, et al.
Published: (2025)
The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
by: Zhu, Yubo, et al.
Published: (2025)
by: Zhu, Yubo, et al.
Published: (2025)
On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models
by: Sancheti, Abhilasha, et al.
Published: (2024)
by: Sancheti, Abhilasha, et al.
Published: (2024)
Adapters for Altering LLM Vocabularies: What Languages Benefit the Most?
by: Han, HyoJung, et al.
Published: (2024)
by: Han, HyoJung, et al.
Published: (2024)
Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style
by: Baumler, Connor, et al.
Published: (2026)
by: Baumler, Connor, et al.
Published: (2026)
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
by: Zhang, Jingshen, et al.
Published: (2024)
by: Zhang, Jingshen, et al.
Published: (2024)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
by: Palta, Shramay, et al.
Published: (2025)
by: Palta, Shramay, et al.
Published: (2025)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
by: Srikanth, Neha, et al.
Published: (2023)
by: Srikanth, Neha, et al.
Published: (2023)
Similar Items
-
Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the U.S
by: Acquaye, Christabel, et al.
Published: (2024) -
Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender?
by: An, Haozhe, et al.
Published: (2024) -
How often are errors in natural language reasoning due to paraphrastic variability?
by: Srikanth, Neha, et al.
Published: (2024) -
Multiple LLM Agents Debate for Equitable Cultural Alignment
by: Ki, Dayeon, et al.
Published: (2025) -
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
by: Palta, Shramay, et al.
Published: (2024)