HyST: LLM-Powered Hybrid Retrieval over Semi-Structured Tabular Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Myung, Jiyoon, Park, Jihyeon, Han, Joohyung
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911120935092224
author Myung, Jiyoon
Park, Jihyeon
Han, Joohyung
author_facet Myung, Jiyoon
Park, Jihyeon
Han, Joohyung
contents User queries in real-world recommendation systems often combine structured constraints (e.g., category, attributes) with unstructured preferences (e.g., product descriptions or reviews). We introduce HyST (Hybrid retrieval over Semi-structured Tabular data), a hybrid retrieval framework that combines LLM-powered structured filtering with semantic embedding search to support complex information needs over semi-structured tabular data. HyST extracts attribute-level constraints from natural language using large language models (LLMs) and applies them as metadata filters, while processing the remaining unstructured query components via embedding-based retrieval. Experiments on a semi-structured benchmark show that HyST consistently outperforms tradtional baselines, highlighting the importance of structured filtering in improving retrieval precision, offering a scalable and accurate solution for real-world user queries.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18048
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HyST: LLM-Powered Hybrid Retrieval over Semi-Structured Tabular Data
Myung, Jiyoon
Park, Jihyeon
Han, Joohyung
Information Retrieval
Artificial Intelligence
User queries in real-world recommendation systems often combine structured constraints (e.g., category, attributes) with unstructured preferences (e.g., product descriptions or reviews). We introduce HyST (Hybrid retrieval over Semi-structured Tabular data), a hybrid retrieval framework that combines LLM-powered structured filtering with semantic embedding search to support complex information needs over semi-structured tabular data. HyST extracts attribute-level constraints from natural language using large language models (LLMs) and applies them as metadata filters, while processing the remaining unstructured query components via embedding-based retrieval. Experiments on a semi-structured benchmark show that HyST consistently outperforms tradtional baselines, highlighting the importance of structured filtering in improving retrieval precision, offering a scalable and accurate solution for real-world user queries.
title HyST: LLM-Powered Hybrid Retrieval over Semi-Structured Tabular Data
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2508.18048