Saved in:
| Main Authors: | Zotos, Leonidas, van Rijn, Hedderik, Nissim, Malvina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.05327 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are You Doubtful? Oh, It Might Be Difficult Then! Exploring the Use of Model Uncertainty for Question Difficulty Estimation
by: Zotos, Leonidas, et al.
Published: (2024)
by: Zotos, Leonidas, et al.
Published: (2024)
The Role of the Availability Heuristic in Multiple-Choice Answering Behaviour
by: Zotos, Leonidas, et al.
Published: (2026)
by: Zotos, Leonidas, et al.
Published: (2026)
NLP Methods May Actually Be Better Than Professors at Estimating Question Difficulty
by: Zotos, Leonidas, et al.
Published: (2025)
by: Zotos, Leonidas, et al.
Published: (2025)
mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models
by: Lai, Huiyuan, et al.
Published: (2024)
by: Lai, Huiyuan, et al.
Published: (2024)
IT5: Text-to-text Pretraining for Italian Language Understanding and Generation
by: Sarti, Gabriele, et al.
Published: (2022)
by: Sarti, Gabriele, et al.
Published: (2022)
Puzzled By ChatGPT? No more! A Jigsaw Puzzle to Promote AI Literacy and Awareness
by: Padovani, Francesca, et al.
Published: (2026)
by: Padovani, Francesca, et al.
Published: (2026)
TACLer: Tailored Curriculum Reinforcement Learning for Efficient Reasoning
by: Lai, Huiyuan, et al.
Published: (2026)
by: Lai, Huiyuan, et al.
Published: (2026)
Multidimensional Consistency Improves Reasoning in Language Models
by: Lai, Huiyuan, et al.
Published: (2025)
by: Lai, Huiyuan, et al.
Published: (2025)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Practising responsibility: Ethics in NLP as a hands-on course
by: Nissim, Malvina, et al.
Published: (2025)
by: Nissim, Malvina, et al.
Published: (2025)
Choosy Babies Need One Coach: Inducing Mode-Seeking Behavior in BabyLlama with Reverse KL Divergence
by: Shi, Shaozhen, et al.
Published: (2024)
by: Shi, Shaozhen, et al.
Published: (2024)
When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue Generation
by: Occhipinti, Daniela, et al.
Published: (2025)
by: Occhipinti, Daniela, et al.
Published: (2025)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
by: Scalena, Daniel, et al.
Published: (2024)
by: Scalena, Daniel, et al.
Published: (2024)
A gentle push funziona benissimo: making instructed models in Italian via contrastive activation steering
by: Scalena, Daniel, et al.
Published: (2024)
by: Scalena, Daniel, et al.
Published: (2024)
Non Verbis, Sed Rebus: Large Language Models are Weak Solvers of Italian Rebuses
by: Sarti, Gabriele, et al.
Published: (2024)
by: Sarti, Gabriele, et al.
Published: (2024)
Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
by: Tomikawa, Yuto, et al.
Published: (2025)
by: Tomikawa, Yuto, et al.
Published: (2025)
UBench: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions
by: Wang, Xunzhi, et al.
Published: (2024)
by: Wang, Xunzhi, et al.
Published: (2024)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
by: Sarti, Gabriele, et al.
Published: (2025)
by: Sarti, Gabriele, et al.
Published: (2025)
Steering Large Language Models for Machine Translation Personalization
by: Scalena, Daniel, et al.
Published: (2025)
by: Scalena, Daniel, et al.
Published: (2025)
Quantifying the Plausibility of Context Reliance in Neural Machine Translation
by: Sarti, Gabriele, et al.
Published: (2023)
by: Sarti, Gabriele, et al.
Published: (2023)
Generating Multiple-Choice Knowledge Questions with Interpretable Difficulty Estimation using Knowledge Graphs and Large Language Models
by: Şakiroğlu, Mehmet Can, et al.
Published: (2026)
by: Şakiroğlu, Mehmet Can, et al.
Published: (2026)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
by: Zhang, Jingshen, et al.
Published: (2024)
by: Zhang, Jingshen, et al.
Published: (2024)
Prediction of Item Difficulty for Reading Comprehension Items by Creation of Annotated Item Repository
by: Kapoor, Radhika, et al.
Published: (2025)
by: Kapoor, Radhika, et al.
Published: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
ARGUS: Seeing the Influence of Narrative Features on Persuasion in Argumentative Texts
by: Nabhani, Sara, et al.
Published: (2026)
by: Nabhani, Sara, et al.
Published: (2026)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
by: Acquaye, Christabel, et al.
Published: (2026)
by: Acquaye, Christabel, et al.
Published: (2026)
QE4PE: Word-level Quality Estimation for Human Post-Editing
by: Sarti, Gabriele, et al.
Published: (2025)
by: Sarti, Gabriele, et al.
Published: (2025)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025)
by: Schmucker, Robin, et al.
Published: (2025)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
by: Deng, Wenqing, et al.
Published: (2024)
by: Deng, Wenqing, et al.
Published: (2024)
Multiple-Choice Questions are Efficient and Robust LLM Evaluators
by: Zhang, Ziyin, et al.
Published: (2024)
by: Zhang, Ziyin, et al.
Published: (2024)
Fine-tuning with HED-IT: The impact of human post-editing for dialogical language models
by: Occhipinti, Daniela, et al.
Published: (2024)
by: Occhipinti, Daniela, et al.
Published: (2024)
Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
by: Park, Yoonah, et al.
Published: (2025)
by: Park, Yoonah, et al.
Published: (2025)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Math Multiple Choice Question Generation via Human-Large Language Model Collaboration
by: Lee, Jaewook, et al.
Published: (2024)
by: Lee, Jaewook, et al.
Published: (2024)
Using Vision + Language Models to Predict Item Difficulty
by: Khan, Samin
Published: (2026)
by: Khan, Samin
Published: (2026)
RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
by: Li, Xinyuan, et al.
Published: (2025)
by: Li, Xinyuan, et al.
Published: (2025)
LLM Distillation for Efficient Few-Shot Multiple Choice Question Answering
by: Sutanto, Patrick, et al.
Published: (2024)
by: Sutanto, Patrick, et al.
Published: (2024)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
by: Lee, Yooseop, et al.
Published: (2025)
by: Lee, Yooseop, et al.
Published: (2025)
Similar Items
-
Are You Doubtful? Oh, It Might Be Difficult Then! Exploring the Use of Model Uncertainty for Question Difficulty Estimation
by: Zotos, Leonidas, et al.
Published: (2024) -
The Role of the Availability Heuristic in Multiple-Choice Answering Behaviour
by: Zotos, Leonidas, et al.
Published: (2026) -
NLP Methods May Actually Be Better Than Professors at Estimating Question Difficulty
by: Zotos, Leonidas, et al.
Published: (2025) -
mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models
by: Lai, Huiyuan, et al.
Published: (2024) -
IT5: Text-to-text Pretraining for Italian Language Understanding and Generation
by: Sarti, Gabriele, et al.
Published: (2022)