ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Heakl, Ahmed, Mohamed, Youssef, Mohamed, Noran, Elsharkawy, Aly, Zaky, Ahmed
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916321868906496
author Heakl, Ahmed
Mohamed, Youssef
Mohamed, Noran
Elsharkawy, Aly
Zaky, Ahmed
author_facet Heakl, Ahmed
Mohamed, Youssef
Mohamed, Noran
Elsharkawy, Aly
Zaky, Ahmed
contents The increasing reliance on online recruitment platforms coupled with the adoption of AI technologies has highlighted the critical need for efficient resume classification methods. However, challenges such as small datasets, lack of standardized resume templates, and privacy concerns hinder the accuracy and effectiveness of existing classification models. In this work, we address these challenges by presenting a comprehensive approach to resume classification. We curated a large-scale dataset of 13,389 resumes from diverse sources and employed Large Language Models (LLMs) such as BERT and Gemma1.1 2B for classification. Our results demonstrate significant improvements over traditional machine learning approaches, with our best model achieving a top-1 accuracy of 92\% and a top-5 accuracy of 97.5\%. These findings underscore the importance of dataset quality and advanced model architectures in enhancing the accuracy and robustness of resume classification systems, thus advancing the field of online recruitment practices.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18125
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models
Heakl, Ahmed
Mohamed, Youssef
Mohamed, Noran
Elsharkawy, Aly
Zaky, Ahmed
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
The increasing reliance on online recruitment platforms coupled with the adoption of AI technologies has highlighted the critical need for efficient resume classification methods. However, challenges such as small datasets, lack of standardized resume templates, and privacy concerns hinder the accuracy and effectiveness of existing classification models. In this work, we address these challenges by presenting a comprehensive approach to resume classification. We curated a large-scale dataset of 13,389 resumes from diverse sources and employed Large Language Models (LLMs) such as BERT and Gemma1.1 2B for classification. Our results demonstrate significant improvements over traditional machine learning approaches, with our best model achieving a top-1 accuracy of 92\% and a top-5 accuracy of 97.5\%. These findings underscore the importance of dataset quality and advanced model architectures in enhancing the accuracy and robustness of resume classification systems, thus advancing the field of online recruitment practices.
title ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2406.18125