A comparative analysis of machine learning algorithms for predicting probabilities of default

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cristescu, Adrian Iulian, Giordano, Matteo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909659158282240
author Cristescu, Adrian Iulian
Giordano, Matteo
author_facet Cristescu, Adrian Iulian
Giordano, Matteo
contents Predicting the probability of default (PD) of prospective loans is a critical objective for financial institutions. In recent years, machine learning (ML) algorithms have achieved remarkable success across a wide variety of prediction tasks; yet, they remain relatively underutilised in credit risk analysis. This paper highlights the opportunities that ML algorithms offer to this field by comparing the performance of five predictive models-Random Forests, Decision Trees, XGBoost, Gradient Boosting and AdaBoost-to the predominantly used logistic regression, over a benchmark dataset from Scheule et al. (Credit Risk Analytics: The R Companion). Our findings underscore the strengths and weaknesses of each method, providing valuable insights into the most effective ML algorithms for PD prediction in the context of loan portfolios.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19789
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A comparative analysis of machine learning algorithms for predicting probabilities of default
Cristescu, Adrian Iulian
Giordano, Matteo
Risk Management
Machine Learning
Applications
Predicting the probability of default (PD) of prospective loans is a critical objective for financial institutions. In recent years, machine learning (ML) algorithms have achieved remarkable success across a wide variety of prediction tasks; yet, they remain relatively underutilised in credit risk analysis. This paper highlights the opportunities that ML algorithms offer to this field by comparing the performance of five predictive models-Random Forests, Decision Trees, XGBoost, Gradient Boosting and AdaBoost-to the predominantly used logistic regression, over a benchmark dataset from Scheule et al. (Credit Risk Analytics: The R Companion). Our findings underscore the strengths and weaknesses of each method, providing valuable insights into the most effective ML algorithms for PD prediction in the context of loan portfolios.
title A comparative analysis of machine learning algorithms for predicting probabilities of default
topic Risk Management
Machine Learning
Applications
url https://arxiv.org/abs/2506.19789