McMining: Automated Discovery of Misconceptions in Student Code
Fuente:
arXiv
Salvato in:
| Autori principali: | , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912640336396288 |
|---|---|
| author | Al-Hossami, Erfan Bunescu, Razvan |
| author_facet | Al-Hossami, Erfan Bunescu, Razvan |
| contents | When learning to code, students often develop misconceptions about various programming language concepts. These can not only lead to bugs or inefficient code, but also slow down the learning of related concepts. In this paper, we introduce McMining, the task of mining programming misconceptions from samples of code from a student. To enable the training and evaluation of McMining systems, we develop an extensible benchmark dataset of misconceptions together with a large set of code samples where these misconceptions are manifested. We then introduce two LLM-based McMiner approaches and through extensive evaluations show that models from the Gemini, Claude, and GPT families are effective at discovering misconceptions in student code. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_08827 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | McMining: Automated Discovery of Misconceptions in Student Code Al-Hossami, Erfan Bunescu, Razvan Software Engineering Artificial Intelligence Computation and Language Computers and Society When learning to code, students often develop misconceptions about various programming language concepts. These can not only lead to bugs or inefficient code, but also slow down the learning of related concepts. In this paper, we introduce McMining, the task of mining programming misconceptions from samples of code from a student. To enable the training and evaluation of McMining systems, we develop an extensible benchmark dataset of misconceptions together with a large set of code samples where these misconceptions are manifested. We then introduce two LLM-based McMiner approaches and through extensive evaluations show that models from the Gemini, Claude, and GPT families are effective at discovering misconceptions in student code. |
| title | McMining: Automated Discovery of Misconceptions in Student Code |
| topic | Software Engineering Artificial Intelligence Computation and Language Computers and Society |
| url | https://arxiv.org/abs/2510.08827 |