ROBUST FEATURE SELECTION IN MULTIVARIATE TIME SERIES CLASSIFICATION: AN EMPIRICAL BENCHMARKING OF GINI, PERMUTATION, AND SHAP
Fuente:
Zenodo
Guardado en:
| Autor principal: | |
|---|---|
| Formato: | Recurso digital |
| Lenguaje: | inglés |
| Publicado: |
Zenodo
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866901901937737728 |
|---|---|
| author | Wuttisasiwat, Nitipat (Ken) |
| author_facet | Wuttisasiwat, Nitipat (Ken) |
| contents | <p class="MsoNormal">Multivariate Time Series Classification presents significant computational challenges due to the need to capture both temporal dependencies and spatial correlations across multiple channels. The Slim-TSF architecture addresses this by utilizing sliding windows to extract statistical features. However, this extraction mechanism substantially expands the feature space, inducing a curse of dimensionality characterized by thousands of highly correlated attributes per instance and resulting in prohibitive computational bottlenecks during the model training phase. The scalability of Slim-TSF therefore relies entirely on the efficiency of its feature selection mechanism.</p> <p class="MsoNormal">This research investigates the computational and predictive trade-offs of substituting Slim-TSF's default embedded feature selection, which uses Gini impurity, with post-hoc evaluation methods, specifically Permutation Importance and SHAP. To conduct this analysis, the Slim-TSF methodology was engineered into a scalable Python library and benchmarked against the 26 standard UEA multivariate datasets.</p> <p class="MsoNormal">The empirical results demonstrate that applying post-hoc feature selection to high-dimensional sliding-window forests introduces severe computational bottlenecks, increasing training times by an order of magnitude, without yielding statistically significant improvements in predictive accuracy. Ultimately, this thesis demonstrates that embedded feature selection provides optimal computational scalability and equivalent predictive representation for production-ready multivariate time series classification architectures.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19588197 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | ROBUST FEATURE SELECTION IN MULTIVARIATE TIME SERIES CLASSIFICATION: AN EMPIRICAL BENCHMARKING OF GINI, PERMUTATION, AND SHAP Wuttisasiwat, Nitipat (Ken) Multivariate Time Series Classification feature importance feature selection <p class="MsoNormal">Multivariate Time Series Classification presents significant computational challenges due to the need to capture both temporal dependencies and spatial correlations across multiple channels. The Slim-TSF architecture addresses this by utilizing sliding windows to extract statistical features. However, this extraction mechanism substantially expands the feature space, inducing a curse of dimensionality characterized by thousands of highly correlated attributes per instance and resulting in prohibitive computational bottlenecks during the model training phase. The scalability of Slim-TSF therefore relies entirely on the efficiency of its feature selection mechanism.</p> <p class="MsoNormal">This research investigates the computational and predictive trade-offs of substituting Slim-TSF's default embedded feature selection, which uses Gini impurity, with post-hoc evaluation methods, specifically Permutation Importance and SHAP. To conduct this analysis, the Slim-TSF methodology was engineered into a scalable Python library and benchmarked against the 26 standard UEA multivariate datasets.</p> <p class="MsoNormal">The empirical results demonstrate that applying post-hoc feature selection to high-dimensional sliding-window forests introduces severe computational bottlenecks, increasing training times by an order of magnitude, without yielding statistically significant improvements in predictive accuracy. Ultimately, this thesis demonstrates that embedded feature selection provides optimal computational scalability and equivalent predictive representation for production-ready multivariate time series classification architectures.</p> |
| title | ROBUST FEATURE SELECTION IN MULTIVARIATE TIME SERIES CLASSIFICATION: AN EMPIRICAL BENCHMARKING OF GINI, PERMUTATION, AND SHAP |
| topic | Multivariate Time Series Classification feature importance feature selection |
| url | https://doi.org/10.5281/zenodo.19588197 |