Ontology-Aligned Embeddings for Data-Driven Labour Market Analytics

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hihn, Heinke, Dittrich, Dennis A. V., Jeske, Carl, Sobral, Cayo Costa, Pais, Helio, Lochmann, Timm
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916935594147840
author Hihn, Heinke
Dittrich, Dennis A. V.
Jeske, Carl
Sobral, Cayo Costa
Pais, Helio
Lochmann, Timm
author_facet Hihn, Heinke
Dittrich, Dennis A. V.
Jeske, Carl
Sobral, Cayo Costa
Pais, Helio
Lochmann, Timm
contents The limited ability to reason across occupational data from different sources is a long-standing bottleneck for data-driven labour market analytics. Previous research has relied on hand-crafted ontologies that allow such reasoning but are computationally expensive and require careful maintenance by human experts. The rise of language processing machine learning models offers a scalable alternative by learning shared semantic spaces that bridge diverse occupational vocabularies without extensive human curation. We present an embedding-based alignment process that links any free-form German job title to two established ontologies - the German Klassifikation der Berufe and the International Standard Classification of Education. Using publicly available data from the German Federal Employment Agency, we construct a dataset to fine-tune a Sentence-BERT model to learn the structure imposed by the ontologies. The enriched pairs (job title, embedding) define a similarity graph structure that we can use for efficient approximate nearest-neighbour search, allowing us to frame the classification process as a semantic search problem. This allows for greater flexibility, e.g., adding more classes. We discuss design decisions, open challenges, and outline ongoing work on extending the graph with other ontologies and multilingual titles.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04942
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ontology-Aligned Embeddings for Data-Driven Labour Market Analytics
Hihn, Heinke
Dittrich, Dennis A. V.
Jeske, Carl
Sobral, Cayo Costa
Pais, Helio
Lochmann, Timm
Machine Learning
The limited ability to reason across occupational data from different sources is a long-standing bottleneck for data-driven labour market analytics. Previous research has relied on hand-crafted ontologies that allow such reasoning but are computationally expensive and require careful maintenance by human experts. The rise of language processing machine learning models offers a scalable alternative by learning shared semantic spaces that bridge diverse occupational vocabularies without extensive human curation. We present an embedding-based alignment process that links any free-form German job title to two established ontologies - the German Klassifikation der Berufe and the International Standard Classification of Education. Using publicly available data from the German Federal Employment Agency, we construct a dataset to fine-tune a Sentence-BERT model to learn the structure imposed by the ontologies. The enriched pairs (job title, embedding) define a similarity graph structure that we can use for efficient approximate nearest-neighbour search, allowing us to frame the classification process as a semantic search problem. This allows for greater flexibility, e.g., adding more classes. We discuss design decisions, open challenges, and outline ongoing work on extending the graph with other ontologies and multilingual titles.
title Ontology-Aligned Embeddings for Data-Driven Labour Market Analytics
topic Machine Learning
url https://arxiv.org/abs/2509.04942