SemBench: A Benchmark for Semantic Query Processing Engines

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lao, Jiale, Zimmerer, Andreas, Ovcharenko, Olga, Cong, Tianji, Russo, Matthew, Vitagliano, Gerardo, Cochez, Michael, Özcan, Fatma, Gupta, Gautam, Hottelier, Thibaud, Jagadish, H. V., Kissel, Kris, Schelter, Sebastian, Kipf, Andreas, Trummer, Immanuel
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910054468288512
author Lao, Jiale
Zimmerer, Andreas
Ovcharenko, Olga
Cong, Tianji
Russo, Matthew
Vitagliano, Gerardo
Cochez, Michael
Özcan, Fatma
Gupta, Gautam
Hottelier, Thibaud
Jagadish, H. V.
Kissel, Kris
Schelter, Sebastian
Kipf, Andreas
Trummer, Immanuel
author_facet Lao, Jiale
Zimmerer, Andreas
Ovcharenko, Olga
Cong, Tianji
Russo, Matthew
Vitagliano, Gerardo
Cochez, Michael
Özcan, Fatma
Gupta, Gautam
Hottelier, Thibaud
Jagadish, H. V.
Kissel, Kris
Schelter, Sebastian
Kipf, Andreas
Trummer, Immanuel
contents We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on generative and reasoning capabilities of state-of-the-art large language models (LLMs). They extend SQL with semantic operators, configured by natural language instructions, that are evaluated via LLMs and enable users to perform various operations on multimodal data. Our benchmark introduces diversity across three key dimensions: scenarios, modalities, and operators. Included are scenarios ranging from movie review analysis to car damage detection. Within these scenarios, we cover different data modalities, including images, audio, and text. Finally, the queries involve a diverse set of operators, including semantic filters, joins, mappings, ranking, and classification operators. We evaluated our benchmark on three academic systems (LOTUS, Palimpzest, and ThalamusDB) and one industrial system, Google BigQuery. Although these results reflect a snapshot of systems under continuous development, our study offers crucial insights into their current strengths and weaknesses, illuminating promising directions for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01716
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SemBench: A Benchmark for Semantic Query Processing Engines
Lao, Jiale
Zimmerer, Andreas
Ovcharenko, Olga
Cong, Tianji
Russo, Matthew
Vitagliano, Gerardo
Cochez, Michael
Özcan, Fatma
Gupta, Gautam
Hottelier, Thibaud
Jagadish, H. V.
Kissel, Kris
Schelter, Sebastian
Kipf, Andreas
Trummer, Immanuel
Databases
Machine Learning
We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on generative and reasoning capabilities of state-of-the-art large language models (LLMs). They extend SQL with semantic operators, configured by natural language instructions, that are evaluated via LLMs and enable users to perform various operations on multimodal data. Our benchmark introduces diversity across three key dimensions: scenarios, modalities, and operators. Included are scenarios ranging from movie review analysis to car damage detection. Within these scenarios, we cover different data modalities, including images, audio, and text. Finally, the queries involve a diverse set of operators, including semantic filters, joins, mappings, ranking, and classification operators. We evaluated our benchmark on three academic systems (LOTUS, Palimpzest, and ThalamusDB) and one industrial system, Google BigQuery. Although these results reflect a snapshot of systems under continuous development, our study offers crucial insights into their current strengths and weaknesses, illuminating promising directions for future research.
title SemBench: A Benchmark for Semantic Query Processing Engines
topic Databases
Machine Learning
url https://arxiv.org/abs/2511.01716