Towards Accurate and Efficient Document Analytics with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Yiming, Hulsebos, Madelon, Ma, Ruiying, Shankar, Shreya, Zeigham, Sepanta, Parameswaran, Aditya G., Wu, Eugene |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Task Cascades for Efficient Unstructured Data Processing
by: Shankar, Shreya, et al.
Published: (2026)
by: Shankar, Shreya, et al.
Published: (2026)
LLM-Powered Proactive Data Systems
by: Zeighami, Sepanta, et al.
Published: (2025)
by: Zeighami, Sepanta, et al.
Published: (2025)
Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees
by: Zeighami, Sepanta, et al.
Published: (2025)
by: Zeighami, Sepanta, et al.
Published: (2025)
Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees
by: Zeighami, Sepanta, et al.
Published: (2025)
by: Zeighami, Sepanta, et al.
Published: (2025)
SPADE: Synthesizing Data Quality Assertions for Large Language Model Pipelines
by: Shankar, Shreya, et al.
Published: (2024)
by: Shankar, Shreya, et al.
Published: (2024)
Semantic Data Processing with Holistic Data Understanding
by: Sun, Youran, et al.
Published: (2026)
by: Sun, Youran, et al.
Published: (2026)
Can AI Agents Answer Your Data Questions? A Benchmark for Data Agents
by: Ma, Ruiying, et al.
Published: (2026)
by: Ma, Ruiying, et al.
Published: (2026)
TARGET: Benchmarking Table Retrieval for Generative Tasks
by: Ji, Xingyu, et al.
Published: (2025)
by: Ji, Xingyu, et al.
Published: (2025)
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
by: Shankar, Shreya, et al.
Published: (2024)
by: Shankar, Shreya, et al.
Published: (2024)
Multi-Objective Agentic Rewrites for Unstructured Data Processing
by: Wei, Lindsey Linxi, et al.
Published: (2025)
by: Wei, Lindsey Linxi, et al.
Published: (2025)
Towards Contextual Sensitive Data Detection
by: Telkamp, Liang, et al.
Published: (2025)
by: Telkamp, Liang, et al.
Published: (2025)
Are We Asking the Right Questions? On Ambiguity in Natural Language Queries for Tabular Data Analysis
by: Gomm, Daniel, et al.
Published: (2025)
by: Gomm, Daniel, et al.
Published: (2025)
Rethinking Dataset Discovery with DataScout
by: Lin, Rachel, et al.
Published: (2025)
by: Lin, Rachel, et al.
Published: (2025)
Arming Data Agents with Tribal Knowledge
by: Agarwal, Shubham, et al.
Published: (2026)
by: Agarwal, Shubham, et al.
Published: (2026)
Steering Semantic Data Processing With DocWrangler
by: Shankar, Shreya, et al.
Published: (2025)
by: Shankar, Shreya, et al.
Published: (2025)
Observatory: Characterizing Embeddings of Relational Tables
by: Cong, Tianji, et al.
Published: (2023)
by: Cong, Tianji, et al.
Published: (2023)
TWIX: Automatically Reconstructing Structured Data from Templatized Documents
by: Lin, Yiming, et al.
Published: (2025)
by: Lin, Yiming, et al.
Published: (2025)
Fine-Grained Table Retrieval Through the Lens of Complex Queries
by: Kosiuk, Wojciech, et al.
Published: (2026)
by: Kosiuk, Wojciech, et al.
Published: (2026)
BiasBuster: a Neural Approach for Accurate Estimation of Population Statistics using Biased Location Data
by: Zeighami, Sepanta, et al.
Published: (2024)
by: Zeighami, Sepanta, et al.
Published: (2024)
Towards Establishing Guaranteed Error for Learned Database Operations
by: Zeighami, Sepanta, et al.
Published: (2024)
by: Zeighami, Sepanta, et al.
Published: (2024)
RAG Without the Lag: Interactive Debugging for Retrieval-Augmented Generation Pipelines
by: Lauro, Quentin Romero, et al.
Published: (2025)
by: Lauro, Quentin Romero, et al.
Published: (2025)
Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
Cocoon: Semantic Table Profiling Using Large Language Models
by: Huang, Zezhou, et al.
Published: (2024)
by: Huang, Zezhou, et al.
Published: (2024)
Data Cleaning Using Large Language Models
by: Zhang, Shuo, et al.
Published: (2024)
by: Zhang, Shuo, et al.
Published: (2024)
Theoretical Analysis of Learned Database Operations under Distribution Shift through Distribution Learnability
by: Zeighami, Sepanta, et al.
Published: (2024)
by: Zeighami, Sepanta, et al.
Published: (2024)
How well do LLMs reason over tabular data, really?
by: Wolff, Cornelius, et al.
Published: (2025)
by: Wolff, Cornelius, et al.
Published: (2025)
SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas
by: Wolff, Cornelius, et al.
Published: (2025)
by: Wolff, Cornelius, et al.
Published: (2025)
MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering
by: Lin, Teng, et al.
Published: (2025)
by: Lin, Teng, et al.
Published: (2025)
PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines
by: Vir, Reya, et al.
Published: (2025)
by: Vir, Reya, et al.
Published: (2025)
NUDGE: Lightweight Non-Parametric Fine-Tuning of Embeddings for Retrieval
by: Zeighami, Sepanta, et al.
Published: (2024)
by: Zeighami, Sepanta, et al.
Published: (2024)
CoddLLM: Empowering Large Language Models for Data Analytics
by: Zhang, Jiani, et al.
Published: (2025)
by: Zhang, Jiani, et al.
Published: (2025)
Access Paths for Efficient Ordering with Large Language Models
by: Zhao, Fuheng, et al.
Published: (2025)
by: Zhao, Fuheng, et al.
Published: (2025)
Towards Autonomous Graph Data Analytics with Analytics-Augmented Generation
by: Wang, Qiange, et al.
Published: (2026)
by: Wang, Qiange, et al.
Published: (2026)
AutoCE: An Accurate and Efficient Model Advisor for Learned Cardinality Estimation
by: Zhang, Jintao, et al.
Published: (2024)
by: Zhang, Jintao, et al.
Published: (2024)
LoPace: A Lossless Optimized Prompt Accurate Compression Engine for Large Language Model Applications
by: Ulla, Aman
Published: (2026)
by: Ulla, Aman
Published: (2026)
Toward Temporal Attribution Analytics in Dataflows
by: Kosyfaki, Chrysanthi, et al.
Published: (2026)
by: Kosyfaki, Chrysanthi, et al.
Published: (2026)
PLOP: Cost-Based Placement of Semantic Operators in Hybrid Query Plans
by: Mang, Qiuyang, et al.
Published: (2026)
by: Mang, Qiuyang, et al.
Published: (2026)
RAC: Relation-Aware Cache Replacement for Large Language Models
by: Wu, Yuchong, et al.
Published: (2026)
by: Wu, Yuchong, et al.
Published: (2026)
EmpireDB: Data System to Accelerate Computational Sciences
by: Alabi, Daniel, et al.
Published: (2024)
by: Alabi, Daniel, et al.
Published: (2024)
DB-GPT-Hub: Towards Open Benchmarking Text-to-SQL Empowered by Large Language Models
by: Zhou, Fan, et al.
Published: (2024)
by: Zhou, Fan, et al.
Published: (2024)
Similar Items
-
Task Cascades for Efficient Unstructured Data Processing
by: Shankar, Shreya, et al.
Published: (2026) -
LLM-Powered Proactive Data Systems
by: Zeighami, Sepanta, et al.
Published: (2025) -
Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees
by: Zeighami, Sepanta, et al.
Published: (2025) -
Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees
by: Zeighami, Sepanta, et al.
Published: (2025) -
SPADE: Synthesizing Data Quality Assertions for Large Language Model Pipelines
by: Shankar, Shreya, et al.
Published: (2024)