DP-Bench: A Benchmark for Evaluating Data Product Creation Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Chowdhury, Faisal, Shirai, Sola, Dash, Sarthak, Mihindukulasooriya, Nandana, Samulowitz, Horst |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation
por: Kashyap, Satyananda, et al.
Publicado: (2025)
por: Kashyap, Satyananda, et al.
Publicado: (2025)
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
por: Zhang, Liangliang, et al.
Publicado: (2025)
por: Zhang, Liangliang, et al.
Publicado: (2025)
Automatic Prompt Engineering with No Task Cues and No Tuning
por: Chowdhury, Faisal, et al.
Publicado: (2026)
por: Chowdhury, Faisal, et al.
Publicado: (2026)
Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks
por: Vogel, Liane, et al.
Publicado: (2026)
por: Vogel, Liane, et al.
Publicado: (2026)
Automatic Prompt Optimization for Knowledge Graph Construction: Insights from an Empirical Study
por: Mihindukulasooriya, Nandana, et al.
Publicado: (2025)
por: Mihindukulasooriya, Nandana, et al.
Publicado: (2025)
SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching
por: Kang, Inwon, et al.
Publicado: (2026)
por: Kang, Inwon, et al.
Publicado: (2026)
BenchPress: A Human-in-the-Loop Annotation System for Rapid Text-to-SQL Benchmark Curation
por: Wenz, Fabian, et al.
Publicado: (2025)
por: Wenz, Fabian, et al.
Publicado: (2025)
Agentic Control Center for Data Product Optimization
por: Tamilselvan, Priyadarshini, et al.
Publicado: (2026)
por: Tamilselvan, Priyadarshini, et al.
Publicado: (2026)
RADAR: Benchmarking Language Models on Imperfect Tabular Data
por: Gu, Ken, et al.
Publicado: (2025)
por: Gu, Ken, et al.
Publicado: (2025)
PersonalHomeBench: Evaluating Agents in Personalized Smart Homes
por: Bharadwaj, Manasa, et al.
Publicado: (2026)
por: Bharadwaj, Manasa, et al.
Publicado: (2026)
PM-LLM-Benchmark: Evaluating Large Language Models on Process Mining Tasks
por: Berti, Alessandro, et al.
Publicado: (2024)
por: Berti, Alessandro, et al.
Publicado: (2024)
AutoDCWorkflow: LLM-based Data Cleaning Workflow Auto-Generation and Benchmark
por: Li, Lan, et al.
Publicado: (2024)
por: Li, Lan, et al.
Publicado: (2024)
Knowledge Base Construction for Knowledge-Augmented Text-to-SQL
por: Baek, Jinheon, et al.
Publicado: (2025)
por: Baek, Jinheon, et al.
Publicado: (2025)
Natural Language Interfaces for Spatial and Temporal Databases: A Comprehensive Overview of Methods, Taxonomy, and Future Directions
por: Acharja, Samya, et al.
Publicado: (2026)
por: Acharja, Samya, et al.
Publicado: (2026)
Evaluating the Data Model Robustness of Text-to-SQL Systems Based on Real User Queries
por: Fürst, Jonathan, et al.
Publicado: (2024)
por: Fürst, Jonathan, et al.
Publicado: (2024)
SQUiD: Synthesizing Relational Databases from Unstructured Text
por: Sadia, Mushtari, et al.
Publicado: (2025)
por: Sadia, Mushtari, et al.
Publicado: (2025)
TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
por: Xie, Jiacheng, et al.
Publicado: (2025)
por: Xie, Jiacheng, et al.
Publicado: (2025)
SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering
por: Lee, Gyubok, et al.
Publicado: (2025)
por: Lee, Gyubok, et al.
Publicado: (2025)
RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
por: Satriani, Dario, et al.
Publicado: (2025)
por: Satriani, Dario, et al.
Publicado: (2025)
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
por: Li, Fangjun, et al.
Publicado: (2024)
por: Li, Fangjun, et al.
Publicado: (2024)
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
por: Tang, Zirui, et al.
Publicado: (2026)
por: Tang, Zirui, et al.
Publicado: (2026)
Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks
por: Yin, Jiaqi, et al.
Publicado: (2025)
por: Yin, Jiaqi, et al.
Publicado: (2025)
MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering
por: Lin, Teng, et al.
Publicado: (2025)
por: Lin, Teng, et al.
Publicado: (2025)
PARROT: A Benchmark for Evaluating LLMs in Cross-System SQL Translation
por: Zhou, Wei, et al.
Publicado: (2025)
por: Zhou, Wei, et al.
Publicado: (2025)
LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs
por: Meyer, Lars-Peter, et al.
Publicado: (2025)
por: Meyer, Lars-Peter, et al.
Publicado: (2025)
Context-Driven Index Trimming: A Data Quality Perspective to Enhancing Precision of RALMs
por: Ma, Kexin, et al.
Publicado: (2024)
por: Ma, Kexin, et al.
Publicado: (2024)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
por: Zhou, Yixi, et al.
Publicado: (2026)
por: Zhou, Yixi, et al.
Publicado: (2026)
Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
por: Cai, Qifeng, et al.
Publicado: (2025)
por: Cai, Qifeng, et al.
Publicado: (2025)
Comprehensive Evaluation for a Large Scale Knowledge Graph Question Answering Service
por: Potdar, Saloni, et al.
Publicado: (2025)
por: Potdar, Saloni, et al.
Publicado: (2025)
The Semantic Ladder: A Framework for Progressive Formalization of Natural Language Content for Knowledge Graphs and AI Systems
por: Vogt, Lars
Publicado: (2026)
por: Vogt, Lars
Publicado: (2026)
Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data
por: Tjuka, Annika, et al.
Publicado: (2025)
por: Tjuka, Annika, et al.
Publicado: (2025)
MAPS: A Multilingual Benchmark for Agent Performance and Security
por: Hofman, Omer, et al.
Publicado: (2025)
por: Hofman, Omer, et al.
Publicado: (2025)
CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era
por: Feng, Yanlin, et al.
Publicado: (2024)
por: Feng, Yanlin, et al.
Publicado: (2024)
OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
Effectiveness of Prompt Optimization in NL2SQL Systems
por: Gurajada, Sairam, et al.
Publicado: (2025)
por: Gurajada, Sairam, et al.
Publicado: (2025)
Grounding Natural Language to SQL Translation with Data-Based Self-Explanations
por: Fan, Yuankai, et al.
Publicado: (2024)
por: Fan, Yuankai, et al.
Publicado: (2024)
Natural Language Querying System Through Entity Enrichment
por: Amavi, Joshua, et al.
Publicado: (2024)
por: Amavi, Joshua, et al.
Publicado: (2024)
LLM and Agent-Driven Data Analysis: A Systematic Approach for Enterprise Applications and System-level Deployment
por: Wang, Xi, et al.
Publicado: (2025)
por: Wang, Xi, et al.
Publicado: (2025)
LLM-R2: A Large Language Model Enhanced Rule-based Rewrite System for Boosting Query Efficiency
por: Li, Zhaodonghui, et al.
Publicado: (2024)
por: Li, Zhaodonghui, et al.
Publicado: (2024)
Anatomy of a Query: W5H Dimensions and FAR Patterns for Text-to-SQL Evaluation
por: Hertzberg, Vicki Stover, et al.
Publicado: (2026)
por: Hertzberg, Vicki Stover, et al.
Publicado: (2026)
Ejemplares similares
-
StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation
por: Kashyap, Satyananda, et al.
Publicado: (2025) -
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
por: Zhang, Liangliang, et al.
Publicado: (2025) -
Automatic Prompt Engineering with No Task Cues and No Tuning
por: Chowdhury, Faisal, et al.
Publicado: (2026) -
Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks
por: Vogel, Liane, et al.
Publicado: (2026) -
Automatic Prompt Optimization for Knowledge Graph Construction: Insights from an Empirical Study
por: Mihindukulasooriya, Nandana, et al.
Publicado: (2025)