FuDoBa: Fusing Document and Knowledge Graph-based Representations with Bayesian Optimisation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koloski, Boshko, Pollak, Senja, Navigli, Roberto, Škrlj, Blaž
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908441740574720
author Koloski, Boshko
Pollak, Senja
Navigli, Roberto
Škrlj, Blaž
author_facet Koloski, Boshko
Pollak, Senja
Navigli, Roberto
Škrlj, Blaž
contents Building on the success of Large Language Models (LLMs), LLM-based representations have dominated the document representation landscape, achieving great performance on the document embedding benchmarks. However, the high-dimensional, computationally expensive embeddings from LLMs tend to be either too generic or inefficient for domain-specific applications. To address these limitations, we introduce FuDoBa a Bayesian optimisation-based method that integrates LLM-based embeddings with domain-specific structured knowledge, sourced both locally and from external repositories like WikiData. This fusion produces low-dimensional, task-relevant representations while reducing training complexity and yielding interpretable early-fusion weights for enhanced classification performance. We demonstrate the effectiveness of our approach on six datasets in two domains, showing that when paired with robust AutoML-based classifiers, our proposed representation learning approach performs on par with, or surpasses, those produced solely by the proprietary LLM-based embedding baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06622
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FuDoBa: Fusing Document and Knowledge Graph-based Representations with Bayesian Optimisation
Koloski, Boshko
Pollak, Senja
Navigli, Roberto
Škrlj, Blaž
Computation and Language
Building on the success of Large Language Models (LLMs), LLM-based representations have dominated the document representation landscape, achieving great performance on the document embedding benchmarks. However, the high-dimensional, computationally expensive embeddings from LLMs tend to be either too generic or inefficient for domain-specific applications. To address these limitations, we introduce FuDoBa a Bayesian optimisation-based method that integrates LLM-based embeddings with domain-specific structured knowledge, sourced both locally and from external repositories like WikiData. This fusion produces low-dimensional, task-relevant representations while reducing training complexity and yielding interpretable early-fusion weights for enhanced classification performance. We demonstrate the effectiveness of our approach on six datasets in two domains, showing that when paired with robust AutoML-based classifiers, our proposed representation learning approach performs on par with, or surpasses, those produced solely by the proprietary LLM-based embedding baselines.
title FuDoBa: Fusing Document and Knowledge Graph-based Representations with Bayesian Optimisation
topic Computation and Language
url https://arxiv.org/abs/2507.06622