Saved in:
Bibliographic Details
Main Authors: Fei, Yulin, Gao, Yuhui, Xian, Xingyuan, Zhang, Xiaojin, Wu, Tao, Chen, Wei
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2412.20613
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929651587219456
author Fei, Yulin
Gao, Yuhui
Xian, Xingyuan
Zhang, Xiaojin
Wu, Tao
Chen, Wei
author_facet Fei, Yulin
Gao, Yuhui
Xian, Xingyuan
Zhang, Xiaojin
Wu, Tao
Chen, Wei
contents With the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character recognition (Video OCR), has become a crucial capability. This paper introduces a novel benchmark designed to evaluate the video OCR performance of multi-modal models in videos. Comprising 1,028 videos and 2,961 question-answer pairs, this benchmark proposes several key challenges through 6 distinct subtasks: (1) Recognition of text content itself and its basic visual attributes, (2)Semantic and Spatial Comprehension of OCR objects in videos (3) Dynamic Motion detection and Temporal Localization. We developed this benchmark using a semi-automated approach that integrates the OCR ability of image LLMs with manual refinement, balancing efficiency, cost, and data quality. Our resource aims to help advance research in video LLMs and underscores the need for improving OCR ability for video LLMs. The benchmark will be released on https://github.com/YuHuiGao/FG-Bench.git.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study
Fei, Yulin
Gao, Yuhui
Xian, Xingyuan
Zhang, Xiaojin
Wu, Tao
Chen, Wei
Computer Vision and Pattern Recognition
With the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character recognition (Video OCR), has become a crucial capability. This paper introduces a novel benchmark designed to evaluate the video OCR performance of multi-modal models in videos. Comprising 1,028 videos and 2,961 question-answer pairs, this benchmark proposes several key challenges through 6 distinct subtasks: (1) Recognition of text content itself and its basic visual attributes, (2)Semantic and Spatial Comprehension of OCR objects in videos (3) Dynamic Motion detection and Temporal Localization. We developed this benchmark using a semi-automated approach that integrates the OCR ability of image LLMs with manual refinement, balancing efficiency, cost, and data quality. Our resource aims to help advance research in video LLMs and underscores the need for improving OCR ability for video LLMs. The benchmark will be released on https://github.com/YuHuiGao/FG-Bench.git.
title Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.20613