Dusty stellar sources classification by implementing machine learning methods based on spectroscopic observations in the Magellanic Clouds

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghaziasgar, Sepideh, Abdollahi, Mahdi, Javadi, Atefeh, van Loon, Jacco, McDonald, Iain, Oliveira, Joana, Masoudnezhad, Amirhossein, Khosroshahi, Habib, Foing, Bernard, Fazel, Fatemeh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908328271020032
author Ghaziasgar, Sepideh
Abdollahi, Mahdi
Javadi, Atefeh
van Loon, Jacco
McDonald, Iain
Oliveira, Joana
Masoudnezhad, Amirhossein
Khosroshahi, Habib
Foing, Bernard
Fazel, Fatemeh
author_facet Ghaziasgar, Sepideh
Abdollahi, Mahdi
Javadi, Atefeh
van Loon, Jacco
McDonald, Iain
Oliveira, Joana
Masoudnezhad, Amirhossein
Khosroshahi, Habib
Foing, Bernard
Fazel, Fatemeh
contents Dusty stellar point sources are a significant stage in stellar evolution and contribute to the metal enrichment of galaxies. These objects can be classified using photometric and spectroscopic observations with color-magnitude diagrams (CMD) and infrared excesses in spectral energy distributions (SED). We employed supervised machine learning spectral classification to categorize dusty stellar sources, including young stellar objects (YSOs) and evolved stars (oxygen- and carbon-rich asymptotic giant branch stars, AGBs), red supergiants (RSGs), and post-AGB (PAGB) stars in the Large and Small Magellanic Clouds, based on spectroscopic labeled data from the Surveying the Agents of Galaxy Evolution (SAGE) project, which used 12 multiwavelength filters and 618 stellar objects. Despite missing values and uncertainties in the SAGE spectral datasets, we achieved accurate classifications. To address small and imbalanced spectral catalogs, we used the Synthetic Minority Oversampling Technique (SMOTE) to generate synthetic data points. Among models applied before and after data augmentation, the Probabilistic Random Forest (PRF), a tuned Random Forest (RF), achieved the highest total accuracy, reaching $\mathbf{89\%}$ based on recall in categorizing dusty stellar sources. Using SMOTE does not improve the best model's accuracy for the CAGB, PAGB, and RSG classes; it remains $\mathbf{100\%}$, $\mathbf{100\%}$, and $\mathbf{88\%}$, respectively, but shows variations for OAGB and YSO classes. We also collected photometric labeled data similar to the training dataset, classifying them using the top four PRF models with over $\mathbf{87\%}$ accuracy. Multiwavelength data from several studies were classified using a consensus model integrating four top models to present common labels as final predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14332
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dusty stellar sources classification by implementing machine learning methods based on spectroscopic observations in the Magellanic Clouds
Ghaziasgar, Sepideh
Abdollahi, Mahdi
Javadi, Atefeh
van Loon, Jacco
McDonald, Iain
Oliveira, Joana
Masoudnezhad, Amirhossein
Khosroshahi, Habib
Foing, Bernard
Fazel, Fatemeh
Astrophysics of Galaxies
Instrumentation and Methods for Astrophysics
Solar and Stellar Astrophysics
Dusty stellar point sources are a significant stage in stellar evolution and contribute to the metal enrichment of galaxies. These objects can be classified using photometric and spectroscopic observations with color-magnitude diagrams (CMD) and infrared excesses in spectral energy distributions (SED). We employed supervised machine learning spectral classification to categorize dusty stellar sources, including young stellar objects (YSOs) and evolved stars (oxygen- and carbon-rich asymptotic giant branch stars, AGBs), red supergiants (RSGs), and post-AGB (PAGB) stars in the Large and Small Magellanic Clouds, based on spectroscopic labeled data from the Surveying the Agents of Galaxy Evolution (SAGE) project, which used 12 multiwavelength filters and 618 stellar objects. Despite missing values and uncertainties in the SAGE spectral datasets, we achieved accurate classifications. To address small and imbalanced spectral catalogs, we used the Synthetic Minority Oversampling Technique (SMOTE) to generate synthetic data points. Among models applied before and after data augmentation, the Probabilistic Random Forest (PRF), a tuned Random Forest (RF), achieved the highest total accuracy, reaching $\mathbf{89\%}$ based on recall in categorizing dusty stellar sources. Using SMOTE does not improve the best model's accuracy for the CAGB, PAGB, and RSG classes; it remains $\mathbf{100\%}$, $\mathbf{100\%}$, and $\mathbf{88\%}$, respectively, but shows variations for OAGB and YSO classes. We also collected photometric labeled data similar to the training dataset, classifying them using the top four PRF models with over $\mathbf{87\%}$ accuracy. Multiwavelength data from several studies were classified using a consensus model integrating four top models to present common labels as final predictions.
title Dusty stellar sources classification by implementing machine learning methods based on spectroscopic observations in the Magellanic Clouds
topic Astrophysics of Galaxies
Instrumentation and Methods for Astrophysics
Solar and Stellar Astrophysics
url https://arxiv.org/abs/2504.14332