HyST: LLM-Powered Hybrid Retrieval over Semi-Structured Tabular Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911120935092224 |
|---|---|
| author | Myung, Jiyoon Park, Jihyeon Han, Joohyung |
| author_facet | Myung, Jiyoon Park, Jihyeon Han, Joohyung |
| contents | User queries in real-world recommendation systems often combine structured constraints (e.g., category, attributes) with unstructured preferences (e.g., product descriptions or reviews). We introduce HyST (Hybrid retrieval over Semi-structured Tabular data), a hybrid retrieval framework that combines LLM-powered structured filtering with semantic embedding search to support complex information needs over semi-structured tabular data. HyST extracts attribute-level constraints from natural language using large language models (LLMs) and applies them as metadata filters, while processing the remaining unstructured query components via embedding-based retrieval. Experiments on a semi-structured benchmark show that HyST consistently outperforms tradtional baselines, highlighting the importance of structured filtering in improving retrieval precision, offering a scalable and accurate solution for real-world user queries. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_18048 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | HyST: LLM-Powered Hybrid Retrieval over Semi-Structured Tabular Data Myung, Jiyoon Park, Jihyeon Han, Joohyung Information Retrieval Artificial Intelligence User queries in real-world recommendation systems often combine structured constraints (e.g., category, attributes) with unstructured preferences (e.g., product descriptions or reviews). We introduce HyST (Hybrid retrieval over Semi-structured Tabular data), a hybrid retrieval framework that combines LLM-powered structured filtering with semantic embedding search to support complex information needs over semi-structured tabular data. HyST extracts attribute-level constraints from natural language using large language models (LLMs) and applies them as metadata filters, while processing the remaining unstructured query components via embedding-based retrieval. Experiments on a semi-structured benchmark show that HyST consistently outperforms tradtional baselines, highlighting the importance of structured filtering in improving retrieval precision, offering a scalable and accurate solution for real-world user queries. |
| title | HyST: LLM-Powered Hybrid Retrieval over Semi-Structured Tabular Data |
| topic | Information Retrieval Artificial Intelligence |
| url | https://arxiv.org/abs/2508.18048 |