DP-Bench: A Benchmark for Evaluating Data Product Creation Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Chowdhury, Faisal, Shirai, Sola, Dash, Sarthak, Mihindukulasooriya, Nandana, Samulowitz, Horst |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation
by: Kashyap, Satyananda, et al.
Published: (2025)
by: Kashyap, Satyananda, et al.
Published: (2025)
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
by: Zhang, Liangliang, et al.
Published: (2025)
by: Zhang, Liangliang, et al.
Published: (2025)
Automatic Prompt Engineering with No Task Cues and No Tuning
by: Chowdhury, Faisal, et al.
Published: (2026)
by: Chowdhury, Faisal, et al.
Published: (2026)
Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks
by: Vogel, Liane, et al.
Published: (2026)
by: Vogel, Liane, et al.
Published: (2026)
Automatic Prompt Optimization for Knowledge Graph Construction: Insights from an Empirical Study
by: Mihindukulasooriya, Nandana, et al.
Published: (2025)
by: Mihindukulasooriya, Nandana, et al.
Published: (2025)
SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching
by: Kang, Inwon, et al.
Published: (2026)
by: Kang, Inwon, et al.
Published: (2026)
BenchPress: A Human-in-the-Loop Annotation System for Rapid Text-to-SQL Benchmark Curation
by: Wenz, Fabian, et al.
Published: (2025)
by: Wenz, Fabian, et al.
Published: (2025)
Agentic Control Center for Data Product Optimization
by: Tamilselvan, Priyadarshini, et al.
Published: (2026)
by: Tamilselvan, Priyadarshini, et al.
Published: (2026)
RADAR: Benchmarking Language Models on Imperfect Tabular Data
by: Gu, Ken, et al.
Published: (2025)
by: Gu, Ken, et al.
Published: (2025)
PersonalHomeBench: Evaluating Agents in Personalized Smart Homes
by: Bharadwaj, Manasa, et al.
Published: (2026)
by: Bharadwaj, Manasa, et al.
Published: (2026)
PM-LLM-Benchmark: Evaluating Large Language Models on Process Mining Tasks
by: Berti, Alessandro, et al.
Published: (2024)
by: Berti, Alessandro, et al.
Published: (2024)
AutoDCWorkflow: LLM-based Data Cleaning Workflow Auto-Generation and Benchmark
by: Li, Lan, et al.
Published: (2024)
by: Li, Lan, et al.
Published: (2024)
Knowledge Base Construction for Knowledge-Augmented Text-to-SQL
by: Baek, Jinheon, et al.
Published: (2025)
by: Baek, Jinheon, et al.
Published: (2025)
Natural Language Interfaces for Spatial and Temporal Databases: A Comprehensive Overview of Methods, Taxonomy, and Future Directions
by: Acharja, Samya, et al.
Published: (2026)
by: Acharja, Samya, et al.
Published: (2026)
Evaluating the Data Model Robustness of Text-to-SQL Systems Based on Real User Queries
by: Fürst, Jonathan, et al.
Published: (2024)
by: Fürst, Jonathan, et al.
Published: (2024)
SQUiD: Synthesizing Relational Databases from Unstructured Text
by: Sadia, Mushtari, et al.
Published: (2025)
by: Sadia, Mushtari, et al.
Published: (2025)
TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
by: Xie, Jiacheng, et al.
Published: (2025)
by: Xie, Jiacheng, et al.
Published: (2025)
SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering
by: Lee, Gyubok, et al.
Published: (2025)
by: Lee, Gyubok, et al.
Published: (2025)
RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
by: Satriani, Dario, et al.
Published: (2025)
by: Satriani, Dario, et al.
Published: (2025)
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
by: Li, Fangjun, et al.
Published: (2024)
by: Li, Fangjun, et al.
Published: (2024)
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
by: Tang, Zirui, et al.
Published: (2026)
by: Tang, Zirui, et al.
Published: (2026)
Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks
by: Yin, Jiaqi, et al.
Published: (2025)
by: Yin, Jiaqi, et al.
Published: (2025)
MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering
by: Lin, Teng, et al.
Published: (2025)
by: Lin, Teng, et al.
Published: (2025)
PARROT: A Benchmark for Evaluating LLMs in Cross-System SQL Translation
by: Zhou, Wei, et al.
Published: (2025)
by: Zhou, Wei, et al.
Published: (2025)
LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs
by: Meyer, Lars-Peter, et al.
Published: (2025)
by: Meyer, Lars-Peter, et al.
Published: (2025)
Context-Driven Index Trimming: A Data Quality Perspective to Enhancing Precision of RALMs
by: Ma, Kexin, et al.
Published: (2024)
by: Ma, Kexin, et al.
Published: (2024)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
by: Zhou, Yixi, et al.
Published: (2026)
by: Zhou, Yixi, et al.
Published: (2026)
Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
by: Cai, Qifeng, et al.
Published: (2025)
by: Cai, Qifeng, et al.
Published: (2025)
Comprehensive Evaluation for a Large Scale Knowledge Graph Question Answering Service
by: Potdar, Saloni, et al.
Published: (2025)
by: Potdar, Saloni, et al.
Published: (2025)
The Semantic Ladder: A Framework for Progressive Formalization of Natural Language Content for Knowledge Graphs and AI Systems
by: Vogt, Lars
Published: (2026)
by: Vogt, Lars
Published: (2026)
Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data
by: Tjuka, Annika, et al.
Published: (2025)
by: Tjuka, Annika, et al.
Published: (2025)
MAPS: A Multilingual Benchmark for Agent Performance and Security
by: Hofman, Omer, et al.
Published: (2025)
by: Hofman, Omer, et al.
Published: (2025)
CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era
by: Feng, Yanlin, et al.
Published: (2024)
by: Feng, Yanlin, et al.
Published: (2024)
OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Effectiveness of Prompt Optimization in NL2SQL Systems
by: Gurajada, Sairam, et al.
Published: (2025)
by: Gurajada, Sairam, et al.
Published: (2025)
Grounding Natural Language to SQL Translation with Data-Based Self-Explanations
by: Fan, Yuankai, et al.
Published: (2024)
by: Fan, Yuankai, et al.
Published: (2024)
Natural Language Querying System Through Entity Enrichment
by: Amavi, Joshua, et al.
Published: (2024)
by: Amavi, Joshua, et al.
Published: (2024)
LLM and Agent-Driven Data Analysis: A Systematic Approach for Enterprise Applications and System-level Deployment
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
LLM-R2: A Large Language Model Enhanced Rule-based Rewrite System for Boosting Query Efficiency
by: Li, Zhaodonghui, et al.
Published: (2024)
by: Li, Zhaodonghui, et al.
Published: (2024)
Anatomy of a Query: W5H Dimensions and FAR Patterns for Text-to-SQL Evaluation
by: Hertzberg, Vicki Stover, et al.
Published: (2026)
by: Hertzberg, Vicki Stover, et al.
Published: (2026)
Similar Items
-
StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation
by: Kashyap, Satyananda, et al.
Published: (2025) -
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
by: Zhang, Liangliang, et al.
Published: (2025) -
Automatic Prompt Engineering with No Task Cues and No Tuning
by: Chowdhury, Faisal, et al.
Published: (2026) -
Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks
by: Vogel, Liane, et al.
Published: (2026) -
Automatic Prompt Optimization for Knowledge Graph Construction: Insights from an Empirical Study
by: Mihindukulasooriya, Nandana, et al.
Published: (2025)