Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Yulong, Liang, Tianyi, Huang, Xinyue, Cui, Erfei, Wang, Guoqing, Guo, Xu, Li, Chenhui, Liu, Gongshen
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909024017973248
author Zhang, Yulong
Liang, Tianyi
Huang, Xinyue
Cui, Erfei
Wang, Guoqing
Guo, Xu
Li, Chenhui
Liu, Gongshen
author_facet Zhang, Yulong
Liang, Tianyi
Huang, Xinyue
Cui, Erfei
Wang, Guoqing
Guo, Xu
Li, Chenhui
Liu, Gongshen
contents Optical Character Recognition (OCR) is fundamental to Vision-Language Models (VLMs) and high-quality data generation for LLM training. Yet, despite progress in average OCR accuracy, state-of-the-art VLMs still struggle with detecting sample-level errors and lack effective unsupervised quality control. We introduce Consensus Entropy (CE), a training-free, model-agnostic metric that estimates output reliability by measuring inter-model agreement entropy. The core insight is that correct predictions converge in output space, while errors diverge. Based on CE, we develop CE-OCR, a lightweight multi-model framework that verifies outputs by ensemble agreement, selects the best outputs, and further improves efficiency through adaptive routing. Experiments demonstrate that CE is robust for quality verification, improving F1 scores by 42.1% over VLM-as-Judge. CE-OCR achieves consistent OCR gains, outperforming self-consistency and single-model baselines at the same cost. Notably, CE requires no training or supervision, enabling plug-and-play integration. Code: https://github.com/Aslan-yulong/consensus-entropy.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11101
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
Zhang, Yulong
Liang, Tianyi
Huang, Xinyue
Cui, Erfei
Wang, Guoqing
Guo, Xu
Li, Chenhui
Liu, Gongshen
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Optical Character Recognition (OCR) is fundamental to Vision-Language Models (VLMs) and high-quality data generation for LLM training. Yet, despite progress in average OCR accuracy, state-of-the-art VLMs still struggle with detecting sample-level errors and lack effective unsupervised quality control. We introduce Consensus Entropy (CE), a training-free, model-agnostic metric that estimates output reliability by measuring inter-model agreement entropy. The core insight is that correct predictions converge in output space, while errors diverge. Based on CE, we develop CE-OCR, a lightweight multi-model framework that verifies outputs by ensemble agreement, selects the best outputs, and further improves efficiency through adaptive routing. Experiments demonstrate that CE is robust for quality verification, improving F1 scores by 42.1% over VLM-as-Judge. CE-OCR achieves consistent OCR gains, outperforming self-consistency and single-model baselines at the same cost. Notably, CE requires no training or supervision, enabling plug-and-play integration. Code: https://github.com/Aslan-yulong/consensus-entropy.
title Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2504.11101