Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Jianyi, Zhou, Ziyin, Ji, Xu, Liu, Shizhao, Zhao, Zhangchi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2512.11323
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917141619408896
author Zhang, Jianyi
Zhou, Ziyin
Ji, Xu
Liu, Shizhao
Zhao, Zhangchi
author_facet Zhang, Jianyi
Zhou, Ziyin
Ji, Xu
Liu, Shizhao
Zhao, Zhangchi
contents Benefiting from strong and efficient multi-modal alignment strategies, Large Visual Language Models (LVLMs) are able to simulate human visual and reasoning capabilities, such as solving CAPTCHAs. However, existing benchmarks based on visual CAPTCHAs still face limitations. Previous studies, when designing benchmarks and datasets, customized them according to their research objectives. Consequently, these benchmarks cannot comprehensively cover all CAPTCHA types. Notably, there is a dearth of dedicated benchmarks for LVLMs. To address this problem, we introduce a novel CAPTCHA benchmark for the first time, named CAPTURE CAPTCHA for Testing Under Real-world Experiments, specifically for LVLMs. Our benchmark encompasses 4 main CAPTCHA types and 25 sub-types from 31 vendors. The diversity enables a multi-dimensional and thorough evaluation of LVLM performance. CAPTURE features extensive class variety, large-scale data, and unique LVLM-tailored labels, filling the gaps in previous research in terms of data comprehensiveness and labeling pertinence. When evaluated by this benchmark, current LVLMs demonstrate poor performance in solving CAPTCHAs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11323
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CAPTURE: A Benchmark and Evaluation for LVLMs in CAPTCHA Resolving
Zhang, Jianyi
Zhou, Ziyin
Ji, Xu
Liu, Shizhao
Zhao, Zhangchi
Artificial Intelligence
Benefiting from strong and efficient multi-modal alignment strategies, Large Visual Language Models (LVLMs) are able to simulate human visual and reasoning capabilities, such as solving CAPTCHAs. However, existing benchmarks based on visual CAPTCHAs still face limitations. Previous studies, when designing benchmarks and datasets, customized them according to their research objectives. Consequently, these benchmarks cannot comprehensively cover all CAPTCHA types. Notably, there is a dearth of dedicated benchmarks for LVLMs. To address this problem, we introduce a novel CAPTCHA benchmark for the first time, named CAPTURE CAPTCHA for Testing Under Real-world Experiments, specifically for LVLMs. Our benchmark encompasses 4 main CAPTCHA types and 25 sub-types from 31 vendors. The diversity enables a multi-dimensional and thorough evaluation of LVLM performance. CAPTURE features extensive class variety, large-scale data, and unique LVLM-tailored labels, filling the gaps in previous research in terms of data comprehensiveness and labeling pertinence. When evaluated by this benchmark, current LVLMs demonstrate poor performance in solving CAPTCHAs.
title CAPTURE: A Benchmark and Evaluation for LVLMs in CAPTCHA Resolving
topic Artificial Intelligence
url https://arxiv.org/abs/2512.11323