SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gondhalekar, Chinmay, Patel, Urjitkumar, Yeh, Fang-Chun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911579380908032
author Gondhalekar, Chinmay
Patel, Urjitkumar
Yeh, Fang-Chun
author_facet Gondhalekar, Chinmay
Patel, Urjitkumar
Yeh, Fang-Chun
contents Accurate question answering over real spreadsheets remains difficult due to multirow headers, merged cells, and unit annotations that disrupt naive chunking, while rigid SQL views fail on files lacking consistent schemas. We present SQuARE, a hybrid retrieval framework with sheet-level, complexity-aware routing. It computes a continuous score based on header depth and merge density, then routes queries either through structure-preserving chunk retrieval or SQL over an automatically constructed relational representation. A lightweight agent supervises retrieval, refinement, or combination of results across both paths when confidence is low. This design maintains header hierarchies, time labels, and units, ensuring that returned values are faithful to the original cells and straightforward to verify. Evaluated on multi-header corporate balance sheets, a heavily merged World Bank workbook, and diverse public datasets, SQuARE consistently surpasses single-strategy baselines and ChatGPT-4o on both retrieval precision and end-to-end answer accuracy while keeping latency predictable. By decoupling retrieval from model choice, the system is compatible with emerging tabular foundation models and offers a practical bridge toward a more robust table understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04292
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
Gondhalekar, Chinmay
Patel, Urjitkumar
Yeh, Fang-Chun
Computation and Language
68T50, 68T20 (Primary) 68T07, 68U15 (Secondary)
I.2.7; I.2.8; H.3.3; H.3.1; I.5.4
Accurate question answering over real spreadsheets remains difficult due to multirow headers, merged cells, and unit annotations that disrupt naive chunking, while rigid SQL views fail on files lacking consistent schemas. We present SQuARE, a hybrid retrieval framework with sheet-level, complexity-aware routing. It computes a continuous score based on header depth and merge density, then routes queries either through structure-preserving chunk retrieval or SQL over an automatically constructed relational representation. A lightweight agent supervises retrieval, refinement, or combination of results across both paths when confidence is low. This design maintains header hierarchies, time labels, and units, ensuring that returned values are faithful to the original cells and straightforward to verify. Evaluated on multi-header corporate balance sheets, a heavily merged World Bank workbook, and diverse public datasets, SQuARE consistently surpasses single-strategy baselines and ChatGPT-4o on both retrieval precision and end-to-end answer accuracy while keeping latency predictable. By decoupling retrieval from model choice, the system is compatible with emerging tabular foundation models and offers a practical bridge toward a more robust table understanding.
title SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
topic Computation and Language
68T50, 68T20 (Primary) 68T07, 68U15 (Secondary)
I.2.7; I.2.8; H.3.3; H.3.1; I.5.4
url https://arxiv.org/abs/2512.04292