Confidence Calibration in Large Language Model-Based Entity Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kamsteeg, Iris, Cardenas-Cartagena, Juan, van Beers, Floris, Holt, Gineke ten, Tashu, Tsegaye Misikir, Valdenegro-Toro, Matias
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911213224460288
author Kamsteeg, Iris
Cardenas-Cartagena, Juan
van Beers, Floris
Holt, Gineke ten
Tashu, Tsegaye Misikir
Valdenegro-Toro, Matias
author_facet Kamsteeg, Iris
Cardenas-Cartagena, Juan
van Beers, Floris
Holt, Gineke ten
Tashu, Tsegaye Misikir
Valdenegro-Toro, Matias
contents This research aims to explore the intersection of Large Language Models and confidence calibration in Entity Matching. To this end, we perform an empirical study to compare baseline RoBERTa confidences for an Entity Matching task against confidences that are calibrated using Temperature Scaling, Monte Carlo Dropout and Ensembles. We use the Abt-Buy, DBLP-ACM, iTunes-Amazon and Company datasets. The findings indicate that the proposed modified RoBERTa model exhibits a slight overconfidence, with Expected Calibration Error scores ranging from 0.0043 to 0.0552 across datasets. We find that this overconfidence can be mitigated using Temperature Scaling, reducing Expected Calibration Error scores by up to 23.83%.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19557
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Confidence Calibration in Large Language Model-Based Entity Matching
Kamsteeg, Iris
Cardenas-Cartagena, Juan
van Beers, Floris
Holt, Gineke ten
Tashu, Tsegaye Misikir
Valdenegro-Toro, Matias
Computation and Language
Machine Learning
This research aims to explore the intersection of Large Language Models and confidence calibration in Entity Matching. To this end, we perform an empirical study to compare baseline RoBERTa confidences for an Entity Matching task against confidences that are calibrated using Temperature Scaling, Monte Carlo Dropout and Ensembles. We use the Abt-Buy, DBLP-ACM, iTunes-Amazon and Company datasets. The findings indicate that the proposed modified RoBERTa model exhibits a slight overconfidence, with Expected Calibration Error scores ranging from 0.0043 to 0.0552 across datasets. We find that this overconfidence can be mitigated using Temperature Scaling, reducing Expected Calibration Error scores by up to 23.83%.
title Confidence Calibration in Large Language Model-Based Entity Matching
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.19557