Towards Cross-Modal Error Detection with Tables and Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ovcharenko, Olga, Schelter, Sebastian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911209298591744
author Ovcharenko, Olga
Schelter, Sebastian
author_facet Ovcharenko, Olga
Schelter, Sebastian
contents Ensuring data quality at scale remains a persistent challenge for large organizations. Despite recent advances, maintaining accurate and consistent data is still complex, especially when dealing with multiple data modalities. Traditional error detection and correction methods tend to focus on a single modality, typically a table, and often miss cross-modal errors that are common in domains like e-Commerce and healthcare, where image, tabular, and text data co-exist. To address this gap, we take an initial step towards cross-modal error detection in tabular data, by benchmarking several methods. Our evaluation spans four datasets and five baseline approaches. Among them, Cleanlab, a label error detection framework, and DataScope, a data valuation method, perform the best when paired with a strong AutoML framework, achieving the highest F1 scores. Our findings indicate that current methods remain limited, particularly when applied to heavy-tailed real-world data, motivating further research in this area.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12383
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Cross-Modal Error Detection with Tables and Images
Ovcharenko, Olga
Schelter, Sebastian
Machine Learning
Ensuring data quality at scale remains a persistent challenge for large organizations. Despite recent advances, maintaining accurate and consistent data is still complex, especially when dealing with multiple data modalities. Traditional error detection and correction methods tend to focus on a single modality, typically a table, and often miss cross-modal errors that are common in domains like e-Commerce and healthcare, where image, tabular, and text data co-exist. To address this gap, we take an initial step towards cross-modal error detection in tabular data, by benchmarking several methods. Our evaluation spans four datasets and five baseline approaches. Among them, Cleanlab, a label error detection framework, and DataScope, a data valuation method, perform the best when paired with a strong AutoML framework, achieving the highest F1 scores. Our findings indicate that current methods remain limited, particularly when applied to heavy-tailed real-world data, motivating further research in this area.
title Towards Cross-Modal Error Detection with Tables and Images
topic Machine Learning
url https://arxiv.org/abs/2510.12383