Real-Time Trustworthiness Scoring for LLM Structured Outputs and Data Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goh, Hui Wen, Mueller, Jonas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912989693607936
author Goh, Hui Wen
Mueller, Jonas
author_facet Goh, Hui Wen
Mueller, Jonas
contents Structured Outputs from current LLMs exhibit sporadic errors, hindering enterprise AI deployment. We present CONSTRUCT, a real-time uncertainty estimator that scores the trustworthiness of LLM Structured Outputs. Lower-scoring outputs are more likely to contain errors, enabling automatic prioritization of limited human review bandwidth. CONSTRUCT additionally scores the trustworthiness of each field within a Structured Output, helping reviewers quickly identify which parts of the output are incorrect. Our method is suitable for any LLM (including black-box LLM APIs without logprobs), does not require labeled training data or custom model deployment, and supports complex Structured Outputs with heterogeneous fields and nested JSON schemas. We also introduce one of the first public LLM Structured Output benchmarks with reliable ground-truth values. Over this four-dataset benchmark, CONSTRUCT detects errors in outputs from various LLMs (including Gemini 3 and GPT-5) with significantly higher precision/recall than existing techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18014
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Real-Time Trustworthiness Scoring for LLM Structured Outputs and Data Extraction
Goh, Hui Wen
Mueller, Jonas
Computation and Language
Machine Learning
Structured Outputs from current LLMs exhibit sporadic errors, hindering enterprise AI deployment. We present CONSTRUCT, a real-time uncertainty estimator that scores the trustworthiness of LLM Structured Outputs. Lower-scoring outputs are more likely to contain errors, enabling automatic prioritization of limited human review bandwidth. CONSTRUCT additionally scores the trustworthiness of each field within a Structured Output, helping reviewers quickly identify which parts of the output are incorrect. Our method is suitable for any LLM (including black-box LLM APIs without logprobs), does not require labeled training data or custom model deployment, and supports complex Structured Outputs with heterogeneous fields and nested JSON schemas. We also introduce one of the first public LLM Structured Output benchmarks with reliable ground-truth values. Over this four-dataset benchmark, CONSTRUCT detects errors in outputs from various LLMs (including Gemini 3 and GPT-5) with significantly higher precision/recall than existing techniques.
title Real-Time Trustworthiness Scoring for LLM Structured Outputs and Data Extraction
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2603.18014