Neurosymbolic Information Extraction from Transactional Documents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hemmer, Arthur, Coustaty, Mickaël, Bartolo, Nicola, Ogier, Jean-Marc
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914192664035328
author Hemmer, Arthur
Coustaty, Mickaël
Bartolo, Nicola
Ogier, Jean-Marc
author_facet Hemmer, Arthur
Coustaty, Mickaël
Bartolo, Nicola
Ogier, Jean-Marc
contents This paper presents a neurosymbolic framework for information extraction from documents, evaluated on transactional documents. We introduce a schema-based approach that integrates symbolic validation methods to enable more effective zero-shot output and knowledge distillation. The methodology uses language models to generate candidate extractions, which are then filtered through syntactic-, task-, and domain-level validation to ensure adherence to domain-specific arithmetic constraints. Our contributions include a comprehensive schema for transactional documents, relabeled datasets, and an approach for generating high-quality labels for knowledge distillation. Experimental results demonstrate significant improvements in $F_1$-scores and accuracy, highlighting the effectiveness of neurosymbolic validation in transactional document processing.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09666
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neurosymbolic Information Extraction from Transactional Documents
Hemmer, Arthur
Coustaty, Mickaël
Bartolo, Nicola
Ogier, Jean-Marc
Computation and Language
This paper presents a neurosymbolic framework for information extraction from documents, evaluated on transactional documents. We introduce a schema-based approach that integrates symbolic validation methods to enable more effective zero-shot output and knowledge distillation. The methodology uses language models to generate candidate extractions, which are then filtered through syntactic-, task-, and domain-level validation to ensure adherence to domain-specific arithmetic constraints. Our contributions include a comprehensive schema for transactional documents, relabeled datasets, and an approach for generating high-quality labels for knowledge distillation. Experimental results demonstrate significant improvements in $F_1$-scores and accuracy, highlighting the effectiveness of neurosymbolic validation in transactional document processing.
title Neurosymbolic Information Extraction from Transactional Documents
topic Computation and Language
url https://arxiv.org/abs/2512.09666