xpSHACL: Explainable SHACL Validation using Retrieval-Augmented Generation and Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Publio, Gustavo Correa, Gayo, José Emilio Labra
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912476351692800
author Publio, Gustavo Correa
Gayo, José Emilio Labra
author_facet Publio, Gustavo Correa
Gayo, José Emilio Labra
contents Shapes Constraint Language (SHACL) is a powerful language for validating RDF data. Given the recent industry attention to Knowledge Graphs (KGs), more users need to validate linked data properly. However, traditional SHACL validation engines often provide terse reports in English that are difficult for non-technical users to interpret and act upon. This paper presents xpSHACL, an explainable SHACL validation system that addresses this issue by combining rule-based justification trees with retrieval-augmented generation (RAG) and large language models (LLMs) to produce detailed, multilanguage, human-readable explanations for constraint violations. A key feature of xpSHACL is its usage of a Violation KG to cache and reuse explanations, improving efficiency and consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08432
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle xpSHACL: Explainable SHACL Validation using Retrieval-Augmented Generation and Large Language Models
Publio, Gustavo Correa
Gayo, José Emilio Labra
Databases
Computation and Language
Shapes Constraint Language (SHACL) is a powerful language for validating RDF data. Given the recent industry attention to Knowledge Graphs (KGs), more users need to validate linked data properly. However, traditional SHACL validation engines often provide terse reports in English that are difficult for non-technical users to interpret and act upon. This paper presents xpSHACL, an explainable SHACL validation system that addresses this issue by combining rule-based justification trees with retrieval-augmented generation (RAG) and large language models (LLMs) to produce detailed, multilanguage, human-readable explanations for constraint violations. A key feature of xpSHACL is its usage of a Violation KG to cache and reuse explanations, improving efficiency and consistency.
title xpSHACL: Explainable SHACL Validation using Retrieval-Augmented Generation and Large Language Models
topic Databases
Computation and Language
url https://arxiv.org/abs/2507.08432