OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Baek, Jinheon, Jeong, Soyeong, Park, Sangwoo, Yeo, Woongyeong, Kang, Minki, Trirat, Patara, Lee, Heejun, Hwang, Sung Ju
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917542744817664
author Baek, Jinheon
Jeong, Soyeong
Park, Sangwoo
Yeo, Woongyeong
Kang, Minki
Trirat, Patara
Lee, Heejun
Hwang, Sung Ju
author_facet Baek, Jinheon
Jeong, Soyeong
Park, Sangwoo
Yeo, Woongyeong
Kang, Minki
Trirat, Patara
Lee, Heejun
Hwang, Sung Ju
contents Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graphs and property graphs. Existing retrievers, however, operate over one source at a time under a fixed query language, leaving the broader landscape of available knowledge fragmented behind incompatible interfaces. A natural attempt at unification would collapse these sources into a shared space, but this erases the structural affordances (such as schemas, ontologies, compositional operators) that give each source its expressive power. Effective retrieval over diverse knowledge, therefore, requires not homogenization but an overarching layer that meets each source on its own terms. To achieve this, we present OmniRetrieval, a framework that takes any natural-language query, identifies appropriate knowledge sources, and dispatches source-native queries to their native execution engines. Across an extensive benchmark spanning 13 datasets and 309 distinct knowledge bases over text, relational, and graph-structured sources, OmniRetrieval exceeds single-source baselines, demonstrating that it can serve as a general-purpose interface to the heterogeneous sources while preserving the structural distinctions that make each source valuable.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29250
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
Baek, Jinheon
Jeong, Soyeong
Park, Sangwoo
Yeo, Woongyeong
Kang, Minki
Trirat, Patara
Lee, Heejun
Hwang, Sung Ju
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graphs and property graphs. Existing retrievers, however, operate over one source at a time under a fixed query language, leaving the broader landscape of available knowledge fragmented behind incompatible interfaces. A natural attempt at unification would collapse these sources into a shared space, but this erases the structural affordances (such as schemas, ontologies, compositional operators) that give each source its expressive power. Effective retrieval over diverse knowledge, therefore, requires not homogenization but an overarching layer that meets each source on its own terms. To achieve this, we present OmniRetrieval, a framework that takes any natural-language query, identifies appropriate knowledge sources, and dispatches source-native queries to their native execution engines. Across an extensive benchmark spanning 13 datasets and 309 distinct knowledge bases over text, relational, and graph-structured sources, OmniRetrieval exceeds single-source baselines, demonstrating that it can serve as a general-purpose interface to the heterogeneous sources while preserving the structural distinctions that make each source valuable.
title OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2605.29250