Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pang, Cong, Yu, Hongtao, Chen, Zixuan, Lu, Lewei, Lou, Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917138085707776
author Pang, Cong
Yu, Hongtao
Chen, Zixuan
Lu, Lewei
Lou, Xin
author_facet Pang, Cong
Yu, Hongtao
Chen, Zixuan
Lu, Lewei
Lou, Xin
contents Large Vision Language Models (LVLMs) have made remarkable progress, enabling sophisticated vision-language interaction and dialogue applications. However, existing benchmarks primarily focus on reasoning tasks, often neglecting fine-grained recognition, which is crucial for practical application scenarios. To address this gap, we introduce the Fine-grained Recognition Open World (FROW) benchmark, designed for detailed evaluation of LVLMs with GPT-4o. On the basis of that, we propose a novel optimization strategy from two perspectives: \textit{data construction} and \textit{training process}, to improve the performance of LVLMs. Our dataset includes mosaic data, which combines multiple short-answer responses, and open-world data, generated from real-world questions and answers using GPT-4o, creating a comprehensive framework for evaluating fine-grained recognition in LVLMs. Experiments show that mosaic data improves category recognition accuracy by 1\% and open-world data boosts FROW benchmark accuracy by 10\%-20\% and content accuracy by 6\%-12\%. Meanwhile, incorporating fine-grained data into the pre-training phase can improve the model's category recognition accuracy by up to 10\%. The benchmark will be available at https://github.com/pc-inno/FROW.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10384
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
Pang, Cong
Yu, Hongtao
Chen, Zixuan
Lu, Lewei
Lou, Xin
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Vision Language Models (LVLMs) have made remarkable progress, enabling sophisticated vision-language interaction and dialogue applications. However, existing benchmarks primarily focus on reasoning tasks, often neglecting fine-grained recognition, which is crucial for practical application scenarios. To address this gap, we introduce the Fine-grained Recognition Open World (FROW) benchmark, designed for detailed evaluation of LVLMs with GPT-4o. On the basis of that, we propose a novel optimization strategy from two perspectives: \textit{data construction} and \textit{training process}, to improve the performance of LVLMs. Our dataset includes mosaic data, which combines multiple short-answer responses, and open-world data, generated from real-world questions and answers using GPT-4o, creating a comprehensive framework for evaluating fine-grained recognition in LVLMs. Experiments show that mosaic data improves category recognition accuracy by 1\% and open-world data boosts FROW benchmark accuracy by 10\%-20\% and content accuracy by 6\%-12\%. Meanwhile, incorporating fine-grained data into the pre-training phase can improve the model's category recognition accuracy by up to 10\%. The benchmark will be available at https://github.com/pc-inno/FROW.
title Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.10384