ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Zanoli, Christopher, Giovannini, Andrea, Jin, Tengjun, Klimovic, Ana, Perlitz, Yotam |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
by: Jin, Tengjun, et al.
Published: (2025)
by: Jin, Tengjun, et al.
Published: (2025)
Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
by: Jin, Tengjun, et al.
Published: (2026)
by: Jin, Tengjun, et al.
Published: (2026)
Mixtera: A Data Plane for Foundation Model Training
by: Böther, Maximilian, et al.
Published: (2025)
by: Böther, Maximilian, et al.
Published: (2025)
From Lossy to Verified: A Provenance-Aware Tiered Memory for Agents
by: Zhu, Qiming, et al.
Published: (2026)
by: Zhu, Qiming, et al.
Published: (2026)
Bootstrapping Learned Cost Models with Synthetic SQL Queries
by: Nidd, Michael, et al.
Published: (2025)
by: Nidd, Michael, et al.
Published: (2025)
DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026)
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026)
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
by: Tang, Zirui, et al.
Published: (2026)
by: Tang, Zirui, et al.
Published: (2026)
KramaBench: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes
by: Lai, Eugenie, et al.
Published: (2025)
by: Lai, Eugenie, et al.
Published: (2025)
LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs
by: Meyer, Lars-Peter, et al.
Published: (2025)
by: Meyer, Lars-Peter, et al.
Published: (2025)
Toward Real-World Table Agents: Capabilities, Workflows, and Design Principles for LLM-based Table Intelligence
by: Tian, Jiaming, et al.
Published: (2025)
by: Tian, Jiaming, et al.
Published: (2025)
RelBench: A Benchmark for Deep Learning on Relational Databases
by: Robinson, Joshua, et al.
Published: (2024)
by: Robinson, Joshua, et al.
Published: (2024)
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
by: Keren, Tomer, et al.
Published: (2026)
by: Keren, Tomer, et al.
Published: (2026)
IndicDB -- Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages
by: Dawar, Aviral, et al.
Published: (2026)
by: Dawar, Aviral, et al.
Published: (2026)
PersonalHomeBench: Evaluating Agents in Personalized Smart Homes
by: Bharadwaj, Manasa, et al.
Published: (2026)
by: Bharadwaj, Manasa, et al.
Published: (2026)
Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory
by: Orogat, Abdelghny, et al.
Published: (2026)
by: Orogat, Abdelghny, et al.
Published: (2026)
VERSA: Verified Event Data Format for Reliable Soccer Analytics
by: Jo, Geonhee, et al.
Published: (2026)
by: Jo, Geonhee, et al.
Published: (2026)
Re-Thinking Process Mining in the AI-Based Agents Era
by: Berti, Alessandro, et al.
Published: (2024)
by: Berti, Alessandro, et al.
Published: (2024)
CMDBench: A Benchmark for Coarse-to-fine Multimodal Data Discovery in Compound AI Systems
by: Feng, Yanlin, et al.
Published: (2024)
by: Feng, Yanlin, et al.
Published: (2024)
Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
Safe, Untrusted, "Proof-Carrying" AI Agents: toward the agentic lakehouse
by: Tagliabue, Jacopo, et al.
Published: (2025)
by: Tagliabue, Jacopo, et al.
Published: (2025)
AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery
by: Kleczek, Darek, et al.
Published: (2026)
by: Kleczek, Darek, et al.
Published: (2026)
AI-Driven Frameworks for Enhancing Data Quality in Big Data Ecosystems: Error_Detection, Correction, and Metadata Integration
by: Elouataoui, Widad
Published: (2024)
by: Elouataoui, Widad
Published: (2024)
Sonar-TS: Search-Then-Verify Natural Language Querying for Time Series Databases
by: Tan, Zhao, et al.
Published: (2026)
by: Tan, Zhao, et al.
Published: (2026)
A Survey of Large Language Model-Based Generative AI for Text-to-SQL: Benchmarks, Applications, Use Cases, and Challenges
by: Singh, Aditi, et al.
Published: (2024)
by: Singh, Aditi, et al.
Published: (2024)
EpiCastBench: Datasets and Benchmarks for Multivariate Epidemic Forecasting
by: Panja, Madhurima, et al.
Published: (2026)
by: Panja, Madhurima, et al.
Published: (2026)
Navigating Tabular Data Synthesis Research: Understanding User Needs and Tool Capabilities
by: R., Maria F. Davila, et al.
Published: (2024)
by: R., Maria F. Davila, et al.
Published: (2024)
PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?
by: Xu, Jingzhe, et al.
Published: (2026)
by: Xu, Jingzhe, et al.
Published: (2026)
Modyn: Data-Centric Machine Learning Pipeline Orchestration
by: Böther, Maximilian, et al.
Published: (2023)
by: Böther, Maximilian, et al.
Published: (2023)
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
by: Watanabe, Yusuke, et al.
Published: (2026)
by: Watanabe, Yusuke, et al.
Published: (2026)
Global Benchmark Database
by: Iser, Ashlin, et al.
Published: (2024)
by: Iser, Ashlin, et al.
Published: (2024)
A Blueprint Architecture of Compound AI Systems for Enterprise
by: Kandogan, Eser, et al.
Published: (2024)
by: Kandogan, Eser, et al.
Published: (2024)
RAG-Driven Data Quality Governance for Enterprise ERP Systems
by: Vedat, Sedat Bin, et al.
Published: (2025)
by: Vedat, Sedat Bin, et al.
Published: (2025)
PIPE-RDF: An LLM-Assisted Pipeline for Enterprise RDF Benchmarking
by: Ranganath, Suraj
Published: (2026)
by: Ranganath, Suraj
Published: (2026)
Arming Data Agents with Tribal Knowledge
by: Agarwal, Shubham, et al.
Published: (2026)
by: Agarwal, Shubham, et al.
Published: (2026)
The FormAI Dataset: Generative AI in Software Security Through the Lens of Formal Verification
by: Tihanyi, Norbert, et al.
Published: (2023)
by: Tihanyi, Norbert, et al.
Published: (2023)
SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
by: Su, Encheng, et al.
Published: (2026)
by: Su, Encheng, et al.
Published: (2026)
A System and Benchmark for LLM-based Q&A on Heterogeneous Data
by: Fokoue, Achille, et al.
Published: (2024)
by: Fokoue, Achille, et al.
Published: (2024)
AI-Driven Research for Databases
by: Cheng, Audrey, et al.
Published: (2026)
by: Cheng, Audrey, et al.
Published: (2026)
LLM/Agent-as-Data-Analyst: A Survey
by: Tang, Zirui, et al.
Published: (2025)
by: Tang, Zirui, et al.
Published: (2025)
Leveraging Knowledge Graphs and LLMs to Support and Monitor Legislative Systems
by: Colombo, Andrea
Published: (2024)
by: Colombo, Andrea
Published: (2024)
Similar Items
-
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
by: Jin, Tengjun, et al.
Published: (2025) -
Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
by: Jin, Tengjun, et al.
Published: (2026) -
Mixtera: A Data Plane for Foundation Model Training
by: Böther, Maximilian, et al.
Published: (2025) -
From Lossy to Verified: A Provenance-Aware Tiered Memory for Agents
by: Zhu, Qiming, et al.
Published: (2026) -
Bootstrapping Learned Cost Models with Synthetic SQL Queries
by: Nidd, Michael, et al.
Published: (2025)