Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Hulingxiao, Li, Geng, Geng, Zijun, Xu, Jinglin, Peng, Yuxin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913765150162944
author He, Hulingxiao
Li, Geng
Geng, Zijun
Xu, Jinglin
Peng, Yuxin
author_facet He, Hulingxiao
Li, Geng
Geng, Zijun
Xu, Jinglin
Peng, Yuxin
contents Multi-modal large language models (MLLMs) have shown remarkable abilities in various visual understanding tasks. However, MLLMs still struggle with fine-grained visual recognition (FGVR), which aims to identify subordinate-level categories from images. This can negatively impact more advanced capabilities of MLLMs, such as object-centric visual question answering and reasoning. In our study, we revisit three quintessential capabilities of MLLMs for FGVR, including object information extraction, category knowledge reserve, object-category alignment, and position of the root cause as a misalignment problem. To address this issue, we present Finedefics, an MLLM that enhances the model's FGVR capability by incorporating informative attribute descriptions of objects into the training phase. We employ contrastive learning on object-attribute pairs and attribute-category pairs simultaneously and use examples from similar but incorrect categories as hard negatives, naturally bringing representations of visual objects and category names closer. Extensive evaluations across multiple popular FGVR datasets demonstrate that Finedefics outperforms existing MLLMs of comparable parameter sizes, showcasing its remarkable efficacy. The code is available at https://github.com/PKU-ICST-MIPL/Finedefics_ICLR2025.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15140
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
He, Hulingxiao
Li, Geng
Geng, Zijun
Xu, Jinglin
Peng, Yuxin
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multi-modal large language models (MLLMs) have shown remarkable abilities in various visual understanding tasks. However, MLLMs still struggle with fine-grained visual recognition (FGVR), which aims to identify subordinate-level categories from images. This can negatively impact more advanced capabilities of MLLMs, such as object-centric visual question answering and reasoning. In our study, we revisit three quintessential capabilities of MLLMs for FGVR, including object information extraction, category knowledge reserve, object-category alignment, and position of the root cause as a misalignment problem. To address this issue, we present Finedefics, an MLLM that enhances the model's FGVR capability by incorporating informative attribute descriptions of objects into the training phase. We employ contrastive learning on object-attribute pairs and attribute-category pairs simultaneously and use examples from similar but incorrect categories as hard negatives, naturally bringing representations of visual objects and category names closer. Extensive evaluations across multiple popular FGVR datasets demonstrate that Finedefics outperforms existing MLLMs of comparable parameter sizes, showcasing its remarkable efficacy. The code is available at https://github.com/PKU-ICST-MIPL/Finedefics_ICLR2025.
title Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2501.15140