BenchPress: A Human-in-the-Loop Annotation System for Rapid Text-to-SQL Benchmark Curation
Fuente:
arXiv
Saved in:
| Main Authors: | Wenz, Fabian, Bouattour, Omar, Yang, Devin, Choi, Justin, Gregg, Cecil, Tatbul, Nesime, Demiralp, Çağatay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BEAVER: An Enterprise Benchmark for Text-to-SQL
by: Chen, Peter Baile, et al.
Published: (2024)
by: Chen, Peter Baile, et al.
Published: (2024)
Mind the Data Gap: Bridging LLMs to Enterprise Data Integration
by: Kayali, Moe, et al.
Published: (2024)
by: Kayali, Moe, et al.
Published: (2024)
Making LLMs Work for Enterprise Data Tasks
by: Demiralp, Çağatay, et al.
Published: (2024)
by: Demiralp, Çağatay, et al.
Published: (2024)
Kairos: Efficient Temporal Graph Analytics on a Single Machine
by: da Trindade, Joana M. F., et al.
Published: (2024)
by: da Trindade, Joana M. F., et al.
Published: (2024)
Humboldt: Metadata-Driven Extensible Data Discovery
by: Bäuerle, Alex, et al.
Published: (2024)
by: Bäuerle, Alex, et al.
Published: (2024)
CascadeServe: Unlocking Model Cascades for Inference Serving
by: Kossmann, Ferdi, et al.
Published: (2024)
by: Kossmann, Ferdi, et al.
Published: (2024)
An Alternate Agentic AI Architecture (It's About the Data)
by: Wenz, Fabian, et al.
Published: (2026)
by: Wenz, Fabian, et al.
Published: (2026)
Understanding the Effects of Noise in Text-to-SQL: An Examination of the BIRD-Bench Benchmark
by: Wretblad, Niklas, et al.
Published: (2024)
by: Wretblad, Niklas, et al.
Published: (2024)
Extract-Transform-Load for Video Streams
by: Kossmann, Ferdinand, et al.
Published: (2023)
by: Kossmann, Ferdinand, et al.
Published: (2023)
What-if Analysis for Business Professionals: Current Practices and Future Opportunities
by: Gathani, Sneha, et al.
Published: (2022)
by: Gathani, Sneha, et al.
Published: (2022)
ReViSQL: Achieving Human-Level Text-to-SQL
by: Zhu, Yuxuan, et al.
Published: (2026)
by: Zhu, Yuxuan, et al.
Published: (2026)
OAG-Bench: A Human-Curated Benchmark for Academic Graph Mining
by: Zhang, Fanjin, et al.
Published: (2024)
by: Zhang, Fanjin, et al.
Published: (2024)
Text-to-SQL Domain Adaptation via Human-LLM Collaborative Data Annotation
by: Tian, Yuan, et al.
Published: (2025)
by: Tian, Yuan, et al.
Published: (2025)
FLEX: Expert-level False-Less EXecution Metric for Reliable Text-to-SQL Benchmark
by: Kim, Heegyu, et al.
Published: (2024)
by: Kim, Heegyu, et al.
Published: (2024)
ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG
by: Atf, Zahra, et al.
Published: (2025)
by: Atf, Zahra, et al.
Published: (2025)
EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records
by: Lee, Gyubok, et al.
Published: (2023)
by: Lee, Gyubok, et al.
Published: (2023)
Improving Demonstration Diversity by Human-Free Fusing for Text-to-SQL
by: Wang, Dingzirui, et al.
Published: (2024)
by: Wang, Dingzirui, et al.
Published: (2024)
Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI
by: Pangakis, Nicholas, et al.
Published: (2024)
by: Pangakis, Nicholas, et al.
Published: (2024)
DesignerlyLoop: Forming Design Intent through Curated Reasoning for Human-LLM Alignment
by: Wang, Anqi, et al.
Published: (2025)
by: Wang, Anqi, et al.
Published: (2025)
Decoupling SQL Query Hardness Parsing for Text-to-SQL
by: Yi, Jiawen, et al.
Published: (2023)
by: Yi, Jiawen, et al.
Published: (2023)
Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
by: Cai, Qifeng, et al.
Published: (2025)
by: Cai, Qifeng, et al.
Published: (2025)
DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph
by: Lee, Jihyung, et al.
Published: (2025)
by: Lee, Jihyung, et al.
Published: (2025)
DTS-SQL: Decomposed Text-to-SQL with Small Large Language Models
by: Pourreza, Mohammadreza, et al.
Published: (2024)
by: Pourreza, Mohammadreza, et al.
Published: (2024)
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
by: Felkner, Virginia K., et al.
Published: (2024)
by: Felkner, Virginia K., et al.
Published: (2024)
LogicCat: A Chain-of-Thought Text-to-SQL Benchmark for Complex Reasoning
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection
by: Byun, Grace, et al.
Published: (2025)
by: Byun, Grace, et al.
Published: (2025)
SLM-SQL: An Exploration of Small Language Models for Text-to-SQL
by: Sheng, Lei, et al.
Published: (2025)
by: Sheng, Lei, et al.
Published: (2025)
STaR-SQL: Self-Taught Reasoner for Text-to-SQL
by: He, Mingqian, et al.
Published: (2025)
by: He, Mingqian, et al.
Published: (2025)
RH-SQL: Refined Schema and Hardness Prompt for Text-to-SQL
by: Yi, Jiawen, et al.
Published: (2024)
by: Yi, Jiawen, et al.
Published: (2024)
A Human-in/on-the-Loop Framework for Accessible Text Generation
by: Moreno, Lourdes, et al.
Published: (2026)
by: Moreno, Lourdes, et al.
Published: (2026)
FINER-SQL: Boosting Small Language Models for Text-to-SQL
by: Hoang, Thanh Dat, et al.
Published: (2026)
by: Hoang, Thanh Dat, et al.
Published: (2026)
EHR-SeqSQL : A Sequential Text-to-SQL Dataset For Interactively Exploring Electronic Health Records
by: Ryu, Jaehee, et al.
Published: (2024)
by: Ryu, Jaehee, et al.
Published: (2024)
PRAXA: A Grammar for What-If Analysis
by: Gathani, Sneha, et al.
Published: (2025)
by: Gathani, Sneha, et al.
Published: (2025)
Human-in-the-Loop Annotation for Image-Based Engagement Estimation: Assessing the Impact of Model Reliability on Annotation Accuracy
by: Subramanya, Sahana Yadnakudige, et al.
Published: (2025)
by: Subramanya, Sahana Yadnakudige, et al.
Published: (2025)
Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
by: Jin, Tengjun, et al.
Published: (2026)
by: Jin, Tengjun, et al.
Published: (2026)
ChatBench: From Static Benchmarks to Human-AI Evaluation
by: Chang, Serina, et al.
Published: (2025)
by: Chang, Serina, et al.
Published: (2025)
LCTG Bench: LLM Controlled Text Generation Benchmark
by: Kurihara, Kentaro, et al.
Published: (2025)
by: Kurihara, Kentaro, et al.
Published: (2025)
Filter & Align: Leveraging Human Knowledge to Curate Image-Text Data
by: Zhang, Lei, et al.
Published: (2023)
by: Zhang, Lei, et al.
Published: (2023)
Bridging Natural Language and Interactive What-If Interfaces via LLM-Generated Declarative Specification
by: Gathani, Sneha, et al.
Published: (2026)
by: Gathani, Sneha, et al.
Published: (2026)
AmbiSQL: Interactive Ambiguity Detection and Resolution for Text-to-SQL
by: Ding, Zhongjun, et al.
Published: (2025)
by: Ding, Zhongjun, et al.
Published: (2025)
Similar Items
-
BEAVER: An Enterprise Benchmark for Text-to-SQL
by: Chen, Peter Baile, et al.
Published: (2024) -
Mind the Data Gap: Bridging LLMs to Enterprise Data Integration
by: Kayali, Moe, et al.
Published: (2024) -
Making LLMs Work for Enterprise Data Tasks
by: Demiralp, Çağatay, et al.
Published: (2024) -
Kairos: Efficient Temporal Graph Analytics on a Single Machine
by: da Trindade, Joana M. F., et al.
Published: (2024) -
Humboldt: Metadata-Driven Extensible Data Discovery
by: Bäuerle, Alex, et al.
Published: (2024)