JobHop: A Large-Scale Dataset of Career Trajectories

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Johary, Iman, Romero, Raphael, Mara, Alexandru C., De Bie, Tijl
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909883754872832
author Johary, Iman
Romero, Raphael
Mara, Alexandru C.
De Bie, Tijl
author_facet Johary, Iman
Romero, Raphael
Mara, Alexandru C.
De Bie, Tijl
contents Understanding labor market dynamics is essential for policymakers, employers, and job seekers. However, comprehensive datasets that capture real-world career trajectories are scarce. In this paper, we introduce JobHop, a large-scale public dataset derived from anonymized resumes provided by VDAB, the public employment service in Flanders, Belgium. Utilizing Large Language Models (LLMs), we process unstructured resume data to extract structured career information, which is then normalized to standardized ESCO occupation codes using a multi-label classification model. This results in a rich dataset of over 1.67 million work experiences, extracted from and grouped into more than 361,000 user resumes and mapped to standardized ESCO occupation codes, offering valuable insights into real-world occupational transitions. This dataset enables diverse applications, such as analyzing labor market mobility, job stability, and the effects of career breaks on occupational transitions. It also supports career path prediction and other data-driven decision-making processes. To illustrate its potential, we explore key dataset characteristics, including job distributions, career breaks, and job transitions, demonstrating its value for advancing labor market research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07653
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JobHop: A Large-Scale Dataset of Career Trajectories
Johary, Iman
Romero, Raphael
Mara, Alexandru C.
De Bie, Tijl
Computation and Language
Understanding labor market dynamics is essential for policymakers, employers, and job seekers. However, comprehensive datasets that capture real-world career trajectories are scarce. In this paper, we introduce JobHop, a large-scale public dataset derived from anonymized resumes provided by VDAB, the public employment service in Flanders, Belgium. Utilizing Large Language Models (LLMs), we process unstructured resume data to extract structured career information, which is then normalized to standardized ESCO occupation codes using a multi-label classification model. This results in a rich dataset of over 1.67 million work experiences, extracted from and grouped into more than 361,000 user resumes and mapped to standardized ESCO occupation codes, offering valuable insights into real-world occupational transitions. This dataset enables diverse applications, such as analyzing labor market mobility, job stability, and the effects of career breaks on occupational transitions. It also supports career path prediction and other data-driven decision-making processes. To illustrate its potential, we explore key dataset characteristics, including job distributions, career breaks, and job transitions, demonstrating its value for advancing labor market research.
title JobHop: A Large-Scale Dataset of Career Trajectories
topic Computation and Language
url https://arxiv.org/abs/2505.07653