MARIC: Multi-Agent Reasoning for Image Classification

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Seo, Wonduk, Yu, Minhyeong, An, Hyunjin, Lee, Seunghyun
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914045593911296
author Seo, Wonduk
Yu, Minhyeong
An, Hyunjin
Lee, Seunghyun
author_facet Seo, Wonduk
Yu, Minhyeong
An, Hyunjin
Lee, Seunghyun
contents Image classification has traditionally relied on parameter-intensive model training, requiring large-scale annotated datasets and extensive fine tuning to achieve competitive performance. While recent vision language models (VLMs) alleviate some of these constraints, they remain limited by their reliance on single pass representations, often failing to capture complementary aspects of visual content. In this paper, we introduce Multi Agent based Reasoning for Image Classification (MARIC), a multi agent framework that reformulates image classification as a collaborative reasoning process. MARIC first utilizes an Outliner Agent to analyze the global theme of the image and generate targeted prompts. Based on these prompts, three Aspect Agents extract fine grained descriptions along distinct visual dimensions. Finally, a Reasoning Agent synthesizes these complementary outputs through integrated reflection step, producing a unified representation for classification. By explicitly decomposing the task into multiple perspectives and encouraging reflective synthesis, MARIC mitigates the shortcomings of both parameter-heavy training and monolithic VLM reasoning. Experiments on 4 diverse image classification benchmark datasets demonstrate that MARIC significantly outperforms baselines, highlighting the effectiveness of multi-agent visual reasoning for robust and interpretable image classification.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14860
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MARIC: Multi-Agent Reasoning for Image Classification
Seo, Wonduk
Yu, Minhyeong
An, Hyunjin
Lee, Seunghyun
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multiagent Systems
Image classification has traditionally relied on parameter-intensive model training, requiring large-scale annotated datasets and extensive fine tuning to achieve competitive performance. While recent vision language models (VLMs) alleviate some of these constraints, they remain limited by their reliance on single pass representations, often failing to capture complementary aspects of visual content. In this paper, we introduce Multi Agent based Reasoning for Image Classification (MARIC), a multi agent framework that reformulates image classification as a collaborative reasoning process. MARIC first utilizes an Outliner Agent to analyze the global theme of the image and generate targeted prompts. Based on these prompts, three Aspect Agents extract fine grained descriptions along distinct visual dimensions. Finally, a Reasoning Agent synthesizes these complementary outputs through integrated reflection step, producing a unified representation for classification. By explicitly decomposing the task into multiple perspectives and encouraging reflective synthesis, MARIC mitigates the shortcomings of both parameter-heavy training and monolithic VLM reasoning. Experiments on 4 diverse image classification benchmark datasets demonstrate that MARIC significantly outperforms baselines, highlighting the effectiveness of multi-agent visual reasoning for robust and interpretable image classification.
title MARIC: Multi-Agent Reasoning for Image Classification
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multiagent Systems
url https://arxiv.org/abs/2509.14860