RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abhyankar, Nikhil, Chaurasia, Purvi, Kabra, Sanchit, Srivastava, Ananya, Gupta, Vivek, Reddy, Chandan K.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909890651357184
author Abhyankar, Nikhil
Chaurasia, Purvi
Kabra, Sanchit
Srivastava, Ananya
Gupta, Vivek
Reddy, Chandan K.
author_facet Abhyankar, Nikhil
Chaurasia, Purvi
Kabra, Sanchit
Srivastava, Ananya
Gupta, Vivek
Reddy, Chandan K.
contents Existing tabular reasoning benchmarks mostly test models on small, uniform tables, underrepresenting the complexity of real-world data and giving an incomplete view of Large Language Models' (LLMs) reasoning abilities. Real tables are long, heterogeneous, and domain-specific, mixing structured fields with free text and requiring multi-hop reasoning across thousands of tokens. To address this gap, we introduce RUST-BENCH, a benchmark of 7966 questions from 2031 real-world tables spanning two domains: i) RB-Science (NSF grant records) and ii) RB-Sports (NBA statistics). Unlike prior work, RUST-BENCH evaluates LLMs jointly across scale, heterogeneity, domain specificity, and reasoning complexity. Experiments with open-source and proprietary models show that LLMs struggle with heterogeneous schemas and complex multi-hop inference, revealing persistent weaknesses in current architectures and prompting strategies. RUST-BENCH establishes a challenging new testbed for advancing tabular reasoning research.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04491
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
Abhyankar, Nikhil
Chaurasia, Purvi
Kabra, Sanchit
Srivastava, Ananya
Gupta, Vivek
Reddy, Chandan K.
Computation and Language
Artificial Intelligence
Databases
Information Retrieval
Machine Learning
Existing tabular reasoning benchmarks mostly test models on small, uniform tables, underrepresenting the complexity of real-world data and giving an incomplete view of Large Language Models' (LLMs) reasoning abilities. Real tables are long, heterogeneous, and domain-specific, mixing structured fields with free text and requiring multi-hop reasoning across thousands of tokens. To address this gap, we introduce RUST-BENCH, a benchmark of 7966 questions from 2031 real-world tables spanning two domains: i) RB-Science (NSF grant records) and ii) RB-Sports (NBA statistics). Unlike prior work, RUST-BENCH evaluates LLMs jointly across scale, heterogeneity, domain specificity, and reasoning complexity. Experiments with open-source and proprietary models show that LLMs struggle with heterogeneous schemas and complex multi-hop inference, revealing persistent weaknesses in current architectures and prompting strategies. RUST-BENCH establishes a challenging new testbed for advancing tabular reasoning research.
title RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
topic Computation and Language
Artificial Intelligence
Databases
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2511.04491