Guardado en:
Detalles Bibliográficos
Autores principales: Meisenbacher, Stephen, Nestorov, Svetlozar, Norlander, Peter
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2510.01470
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911188689879040
author Meisenbacher, Stephen
Nestorov, Svetlozar
Norlander, Peter
author_facet Meisenbacher, Stephen
Nestorov, Svetlozar
Norlander, Peter
contents Data from online job postings are difficult to access and are not built in a standard or transparent manner. Data included in the standard taxonomy and occupational information database (O*NET) are updated infrequently and based on small survey samples. We adopt O*NET as a framework for building natural language processing tools that extract structured information from job postings. We publish the Job Ad Analysis Toolkit (JAAT), a collection of open-source tools built for this purpose, and demonstrate its reliability and accuracy in out-of-sample and LLM-as-a-Judge testing. We extract more than 10 billion data points from more than 155 million online job ads provided by the National Labor Exchange (NLx) Research Hub, including O*NET tasks, occupation codes, tools, and technologies, as well as wages, skills, industry, and more features. We describe the construction of a dataset of occupation, state, and industry level features aggregated by monthly active jobs from 2015 - 2025. We illustrate the potential for research and future uses in education and workforce development.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01470
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Extracting O*NET Features from the NLx Corpus to Build Public Use Aggregate Labor Market Data
Meisenbacher, Stephen
Nestorov, Svetlozar
Norlander, Peter
Computers and Society
Computation and Language
Data from online job postings are difficult to access and are not built in a standard or transparent manner. Data included in the standard taxonomy and occupational information database (O*NET) are updated infrequently and based on small survey samples. We adopt O*NET as a framework for building natural language processing tools that extract structured information from job postings. We publish the Job Ad Analysis Toolkit (JAAT), a collection of open-source tools built for this purpose, and demonstrate its reliability and accuracy in out-of-sample and LLM-as-a-Judge testing. We extract more than 10 billion data points from more than 155 million online job ads provided by the National Labor Exchange (NLx) Research Hub, including O*NET tasks, occupation codes, tools, and technologies, as well as wages, skills, industry, and more features. We describe the construction of a dataset of occupation, state, and industry level features aggregated by monthly active jobs from 2015 - 2025. We illustrate the potential for research and future uses in education and workforce development.
title Extracting O*NET Features from the NLx Corpus to Build Public Use Aggregate Labor Market Data
topic Computers and Society
Computation and Language
url https://arxiv.org/abs/2510.01470