Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tiwari, Aman, Malay, Shiva Krishna Reddy, Yadav, Vikas, Hashemi, Masoud, Madhusudhan, Sathwik Tejaswi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917901367246848
author Tiwari, Aman
Malay, Shiva Krishna Reddy
Yadav, Vikas
Hashemi, Masoud
Madhusudhan, Sathwik Tejaswi
author_facet Tiwari, Aman
Malay, Shiva Krishna Reddy
Yadav, Vikas
Hashemi, Masoud
Madhusudhan, Sathwik Tejaswi
contents Graph databases like Neo4j are gaining popularity for handling complex, interconnected data, over traditional relational databases in modeling and querying relationships. While translating natural language into SQL queries is well-researched, generating Cypher queries for Neo4j remains relatively underexplored. In this work, we present an automated, LLM-Supervised, pipeline to generate high-quality synthetic data for Text2Cypher. Our Cypher data generation pipeline introduces LLM-As-Database-Filler, a novel strategy for ensuring Cypher query correctness, thus resulting in high quality generations. Using our pipeline, we generate high quality Text2Cypher data - SynthCypher containing 29.8k instances across various domains and queries with varying complexities. Training open-source LLMs like LLaMa-3.1-8B, Mistral-7B, and QWEN-7B on SynthCypher results in performance gains of up to 40% on the Text2Cypher test split and 30% on the SPIDER benchmark, adapted for graph databases.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12612
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
Tiwari, Aman
Malay, Shiva Krishna Reddy
Yadav, Vikas
Hashemi, Masoud
Madhusudhan, Sathwik Tejaswi
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
Graph databases like Neo4j are gaining popularity for handling complex, interconnected data, over traditional relational databases in modeling and querying relationships. While translating natural language into SQL queries is well-researched, generating Cypher queries for Neo4j remains relatively underexplored. In this work, we present an automated, LLM-Supervised, pipeline to generate high-quality synthetic data for Text2Cypher. Our Cypher data generation pipeline introduces LLM-As-Database-Filler, a novel strategy for ensuring Cypher query correctness, thus resulting in high quality generations. Using our pipeline, we generate high quality Text2Cypher data - SynthCypher containing 29.8k instances across various domains and queries with varying complexities. Training open-source LLMs like LLaMa-3.1-8B, Mistral-7B, and QWEN-7B on SynthCypher results in performance gains of up to 40% on the Text2Cypher test split and 30% on the SPIDER benchmark, adapted for graph databases.
title Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2412.12612