UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rogoz, Ana-Cristina, Ionescu, Radu Tudor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PoPreRo: A New Dataset for Popularity Prediction of Romanian Reddit Posts
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2024)
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2024)
Large Multimodal Models for Low-Resource Languages: A Survey
von: Lupascu, Marian, et al.
Veröffentlicht: (2025)
von: Lupascu, Marian, et al.
Veröffentlicht: (2025)
A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2025)
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2025)
Every Character Counts: From Vulnerability to Defense in Phishing Detection
von: Chiper, Maria, et al.
Veröffentlicht: (2025)
von: Chiper, Maria, et al.
Veröffentlicht: (2025)
CLewR: Curriculum Learning with Restarts for Machine Translation Preference Learning
von: Dragomir, Alexandra, et al.
Veröffentlicht: (2026)
von: Dragomir, Alexandra, et al.
Veröffentlicht: (2026)
A Novel Cartography-Based Curriculum Learning Method Applied on RoNLI: The First Romanian Natural Language Inference Corpus
von: Poesina, Eduard, et al.
Veröffentlicht: (2024)
von: Poesina, Eduard, et al.
Veröffentlicht: (2024)
Task-Informed Anti-Curriculum by Masking Improves Downstream Performance on Text
von: Jarca, Andrei, et al.
Veröffentlicht: (2025)
von: Jarca, Andrei, et al.
Veröffentlicht: (2025)
Learning Using Generated Privileged Information by Text-to-Image Diffusion Models
von: Menadil, Rafael-Edy, et al.
Veröffentlicht: (2023)
von: Menadil, Rafael-Edy, et al.
Veröffentlicht: (2023)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
Generating Multiple-Choice Knowledge Questions with Interpretable Difficulty Estimation using Knowledge Graphs and Large Language Models
von: Şakiroğlu, Mehmet Can, et al.
Veröffentlicht: (2026)
von: Şakiroğlu, Mehmet Can, et al.
Veröffentlicht: (2026)
Text Classification Under Class Distribution Shift: A Survey
von: Costache, Adriana Valentina, et al.
Veröffentlicht: (2025)
von: Costache, Adriana Valentina, et al.
Veröffentlicht: (2025)
Machine Unlearning in the Era of Quantum Machine Learning: An Empirical Study
von: Crivoi, Carla, et al.
Veröffentlicht: (2025)
von: Crivoi, Carla, et al.
Veröffentlicht: (2025)
PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction
von: Poesina, Eduard, et al.
Veröffentlicht: (2024)
von: Poesina, Eduard, et al.
Veröffentlicht: (2024)
ExDDV: A New Dataset for Explainable Deepfake Detection in Video
von: Hondru, Vlad, et al.
Veröffentlicht: (2025)
von: Hondru, Vlad, et al.
Veröffentlicht: (2025)
Biomedical Entity Linking as Multiple Choice Question Answering
von: Lin, Zhenxi, et al.
Veröffentlicht: (2024)
von: Lin, Zhenxi, et al.
Veröffentlicht: (2024)
Option-ID Based Elimination For Multiple Choice Questions
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025)
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025)
MOSLD-Bench: Multilingual Open-Set Learning and Discovery Benchmark for Text Categorization
von: Costache, Adriana-Valentina, et al.
Veröffentlicht: (2026)
von: Costache, Adriana-Valentina, et al.
Veröffentlicht: (2026)
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)
Pattern Recognition or Medical Knowledge? The Problem with Multiple-Choice Questions in Medicine
von: Griot, Maxime, et al.
Veröffentlicht: (2024)
von: Griot, Maxime, et al.
Veröffentlicht: (2024)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
A Study on Large Language Models' Limitations in Multiple-Choice Question Answering
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
Multi-Level Feature Distillation of Joint Teachers Trained on Distinct Image Datasets
von: Iordache, Adrian, et al.
Veröffentlicht: (2024)
von: Iordache, Adrian, et al.
Veröffentlicht: (2024)
Curriculum Multi-Task Self-Supervision Improves Lightweight Architectures for Onboard Satellite Hyperspectral Image Segmentation
von: Carlesso, Hugo, et al.
Veröffentlicht: (2025)
von: Carlesso, Hugo, et al.
Veröffentlicht: (2025)
JumpLoRA: Sparse Adapters for Continual Learning in Large Language Models
von: Dragomir, Alexandra, et al.
Veröffentlicht: (2026)
von: Dragomir, Alexandra, et al.
Veröffentlicht: (2026)
AutoMalDesc: Large-Scale Script Analysis for Cyber Threat Research
von: Apostu, Alexandru-Mihai, et al.
Veröffentlicht: (2025)
von: Apostu, Alexandru-Mihai, et al.
Veröffentlicht: (2025)
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
CBM: Curriculum by Masking
von: Jarca, Andrei, et al.
Veröffentlicht: (2024)
von: Jarca, Andrei, et al.
Veröffentlicht: (2024)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2025)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2025)
XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark
von: Ciobanu, Ioan-Paul, et al.
Veröffentlicht: (2025)
von: Ciobanu, Ioan-Paul, et al.
Veröffentlicht: (2025)
Exploring Design Choices for Building Language-Specific LLMs
von: Tejaswi, Atula, et al.
Veröffentlicht: (2024)
von: Tejaswi, Atula, et al.
Veröffentlicht: (2024)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2026)
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2026)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
von: Letzelter, Victor, et al.
Veröffentlicht: (2025)
von: Letzelter, Victor, et al.
Veröffentlicht: (2025)
Harnessing the Power of Semi-Structured Knowledge and LLMs with Triplet-Based Prefiltering for Question Answering
von: Boer, Derian, et al.
Veröffentlicht: (2024)
von: Boer, Derian, et al.
Veröffentlicht: (2024)
Diffusion Models in Vision: A Survey
von: Croitoru, Florinel-Alin, et al.
Veröffentlicht: (2022)
von: Croitoru, Florinel-Alin, et al.
Veröffentlicht: (2022)
3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selection
von: Ding, Hongxin, et al.
Veröffentlicht: (2024)
von: Ding, Hongxin, et al.
Veröffentlicht: (2024)
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PoPreRo: A New Dataset for Popularity Prediction of Romanian Reddit Posts
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2024) -
Large Multimodal Models for Low-Resource Languages: A Survey
von: Lupascu, Marian, et al.
Veröffentlicht: (2025) -
A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2025) -
Every Character Counts: From Vulnerability to Defense in Phishing Detection
von: Chiper, Maria, et al.
Veröffentlicht: (2025) -
CLewR: Curriculum Learning with Restarts for Machine Translation Preference Learning
von: Dragomir, Alexandra, et al.
Veröffentlicht: (2026)