AI-Driven Research for Databases

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheng, Audrey, Ng, Harald, Kabcenell, Aaron, Bailis, Peter, Zaharia, Matei, Ma, Lin, Shi, Xiao, Stoica, Ion
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917389871874048
author Cheng, Audrey
Ng, Harald
Kabcenell, Aaron
Bailis, Peter
Zaharia, Matei
Ma, Lin
Shi, Xiao
Stoica, Ion
author_facet Cheng, Audrey
Ng, Harald
Kabcenell, Aaron
Bailis, Peter
Zaharia, Matei
Ma, Lin
Shi, Xiao
Stoica, Ion
contents As the complexity of modern workloads and hardware increasingly outpaces human research and engineering capacity, existing methods for database performance optimization struggle to keep pace. To address this gap, a new class of techniques, termed AI-Driven Research for Systems (ADRS), uses large language models to automate solution discovery. This approach shifts optimization from manual system design to automated code generation. The key obstacle, however, in applying ADRS is the evaluation pipeline. Since these frameworks rapidly generate hundreds of candidates without human supervision, they depend on fast and accurate feedback from evaluators to converge on effective solutions. Building such evaluators is especially difficult for complex database systems. To enable the practical application of ADRS in this domain, we propose automating the design of evaluators by co-evolving them with the solutions. We demonstrate the effectiveness of this approach through three case studies optimizing buffer management, query rewriting, and index selection. Our automated evaluators enable the discovery of novel algorithms that outperform state-of-the-art baselines (e.g., a deterministic query rewrite policy that achieves up to 6.8x lower latency), demonstrating that addressing the evaluation bottleneck unlocks the potential of ADRS to generate highly optimized, deployable code for next-generation data systems.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06566
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AI-Driven Research for Databases
Cheng, Audrey
Ng, Harald
Kabcenell, Aaron
Bailis, Peter
Zaharia, Matei
Ma, Lin
Shi, Xiao
Stoica, Ion
Databases
Artificial Intelligence
As the complexity of modern workloads and hardware increasingly outpaces human research and engineering capacity, existing methods for database performance optimization struggle to keep pace. To address this gap, a new class of techniques, termed AI-Driven Research for Systems (ADRS), uses large language models to automate solution discovery. This approach shifts optimization from manual system design to automated code generation. The key obstacle, however, in applying ADRS is the evaluation pipeline. Since these frameworks rapidly generate hundreds of candidates without human supervision, they depend on fast and accurate feedback from evaluators to converge on effective solutions. Building such evaluators is especially difficult for complex database systems. To enable the practical application of ADRS in this domain, we propose automating the design of evaluators by co-evolving them with the solutions. We demonstrate the effectiveness of this approach through three case studies optimizing buffer management, query rewriting, and index selection. Our automated evaluators enable the discovery of novel algorithms that outperform state-of-the-art baselines (e.g., a deterministic query rewrite policy that achieves up to 6.8x lower latency), demonstrating that addressing the evaluation bottleneck unlocks the potential of ADRS to generate highly optimized, deployable code for next-generation data systems.
title AI-Driven Research for Databases
topic Databases
Artificial Intelligence
url https://arxiv.org/abs/2604.06566