Identifying Optimal Regression Models For DEM Simulation Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jenkins, B. D., Nicusan, A. L., Neveu, A., Lumay, G., Francqui, F., Seville, J. P. K., Windows-Yule, C. R. K.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916929702199296
author Jenkins, B. D.
Nicusan, A. L.
Neveu, A.
Lumay, G.
Francqui, F.
Seville, J. P. K.
Windows-Yule, C. R. K.
author_facet Jenkins, B. D.
Nicusan, A. L.
Neveu, A.
Lumay, G.
Francqui, F.
Seville, J. P. K.
Windows-Yule, C. R. K.
contents Developing fast regression models (surrogate/metamodels) from DEM data is key for practical industrial application to allow real-time evaluations. However, benchmarking different models is often overlooked in particle technology for regression tasks, as model selection is frequently not the primary research focus. This can lead to the use of suboptimal models, resulting in subpar predictive accuracy, slow evaluations, or poor generalisation, hindering effective real-time decision-making and process optimisation. In this work, we discuss applying k-fold cross-validation to assess regression models for tabular DEM datasets and propose a simple framework for readers to follow to find the optimal model for their data. An example demonstrates its application to a DEM dataset of packing fractions measured in a simple measuring beaker with varying inter-particle properties, namely, average particle diameter, coefficient of restitution, coefficient of sliding friction, coefficient of rolling resistance, and cohesive energy density. Out of 16 different models tested, a histogram-based gradient boosting model was found to be optimal, providing a good fit with acceptable training and inference times.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05308
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Identifying Optimal Regression Models For DEM Simulation Datasets
Jenkins, B. D.
Nicusan, A. L.
Neveu, A.
Lumay, G.
Francqui, F.
Seville, J. P. K.
Windows-Yule, C. R. K.
Computational Physics
Developing fast regression models (surrogate/metamodels) from DEM data is key for practical industrial application to allow real-time evaluations. However, benchmarking different models is often overlooked in particle technology for regression tasks, as model selection is frequently not the primary research focus. This can lead to the use of suboptimal models, resulting in subpar predictive accuracy, slow evaluations, or poor generalisation, hindering effective real-time decision-making and process optimisation. In this work, we discuss applying k-fold cross-validation to assess regression models for tabular DEM datasets and propose a simple framework for readers to follow to find the optimal model for their data. An example demonstrates its application to a DEM dataset of packing fractions measured in a simple measuring beaker with varying inter-particle properties, namely, average particle diameter, coefficient of restitution, coefficient of sliding friction, coefficient of rolling resistance, and cohesive energy density. Out of 16 different models tested, a histogram-based gradient boosting model was found to be optimal, providing a good fit with acceptable training and inference times.
title Identifying Optimal Regression Models For DEM Simulation Datasets
topic Computational Physics
url https://arxiv.org/abs/2508.05308