Predicting First Year Dropout from Pre Enrolment Motivation Statements Using Text Mining

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Soppe, K. F. B., Bagheri, A., Nadi, S., Klugkist, I. G., Wubbels, T., Meij, L. D. N. V. Wijngaards-De
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908548978442240
author Soppe, K. F. B.
Bagheri, A.
Nadi, S.
Klugkist, I. G.
Wubbels, T.
Meij, L. D. N. V. Wijngaards-De
author_facet Soppe, K. F. B.
Bagheri, A.
Nadi, S.
Klugkist, I. G.
Wubbels, T.
Meij, L. D. N. V. Wijngaards-De
contents Preventing student dropout is a major challenge in higher education and it is difficult to predict prior to enrolment which students are likely to drop out and which students are likely to succeed. High School GPA is a strong predictor of dropout, but much variance in dropout remains to be explained. This study focused on predicting university dropout by using text mining techniques with the aim of exhuming information contained in motivation statements written by students. By combining text data with classic predictors of dropout in the form of student characteristics, we attempt to enhance the available set of predictive student characteristics. Our dataset consisted of 7,060 motivation statements of students enrolling in a non-selective bachelor at a Dutch university in 2014 and 2015. Support Vector Machines were trained on 75 percent of the data and several models were estimated on the test data. We used various combinations of student characteristics and text, such as TFiDF, topic modelling, LIWC dictionary. Results showed that, although the combination of text and student characteristics did not improve the prediction of dropout, text analysis alone predicted dropout similarly well as a set of student characteristics. Suggestions for future research are provided.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Predicting First Year Dropout from Pre Enrolment Motivation Statements Using Text Mining
Soppe, K. F. B.
Bagheri, A.
Nadi, S.
Klugkist, I. G.
Wubbels, T.
Meij, L. D. N. V. Wijngaards-De
Computers and Society
Computation and Language
Machine Learning
Applications
Preventing student dropout is a major challenge in higher education and it is difficult to predict prior to enrolment which students are likely to drop out and which students are likely to succeed. High School GPA is a strong predictor of dropout, but much variance in dropout remains to be explained. This study focused on predicting university dropout by using text mining techniques with the aim of exhuming information contained in motivation statements written by students. By combining text data with classic predictors of dropout in the form of student characteristics, we attempt to enhance the available set of predictive student characteristics. Our dataset consisted of 7,060 motivation statements of students enrolling in a non-selective bachelor at a Dutch university in 2014 and 2015. Support Vector Machines were trained on 75 percent of the data and several models were estimated on the test data. We used various combinations of student characteristics and text, such as TFiDF, topic modelling, LIWC dictionary. Results showed that, although the combination of text and student characteristics did not improve the prediction of dropout, text analysis alone predicted dropout similarly well as a set of student characteristics. Suggestions for future research are provided.
title Predicting First Year Dropout from Pre Enrolment Motivation Statements Using Text Mining
topic Computers and Society
Computation and Language
Machine Learning
Applications
url https://arxiv.org/abs/2509.16224