MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yüksel, Arda, Thiem, Gabriel, Walter, Susanne, Felka, Patrick, Werb, Gabriela Alves, Habernal, Ivan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913020662251520
author Yüksel, Arda
Thiem, Gabriel
Walter, Susanne
Felka, Patrick
Werb, Gabriela Alves
Habernal, Ivan
author_facet Yüksel, Arda
Thiem, Gabriel
Walter, Susanne
Felka, Patrick
Werb, Gabriela Alves
Habernal, Ivan
contents Industry classification schemes are integral parts of public and corporate databases as they classify businesses based on economic activity. Due to the size of the company registers, manual annotation is costly, and fine-tuning models with every update in industry classification schemes requires significant data collection. We replicate the manual expert verification by using existing or easily retrievable multimodal resources for industry classification. We present MONETA, the first multimodal industry classification benchmark with text (Website, Wikipedia, Wikidata) and geospatial sources (OpenStreetMap and satellite imagery). Our dataset enlists 1,000 businesses in Europe with 20 economic activity labels according to EU guidelines (NACE). Our training-free baseline reaches 62.10% and 74.10% with open and closed-source Multimodal Large Language Models (MLLM). We observe an increase of up to 22.80% with the combination of multi-turn design, context enrichment, and classification explanations. We will release our dataset and the enhanced guidelines.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07956
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems
Yüksel, Arda
Thiem, Gabriel
Walter, Susanne
Felka, Patrick
Werb, Gabriela Alves
Habernal, Ivan
Artificial Intelligence
Industry classification schemes are integral parts of public and corporate databases as they classify businesses based on economic activity. Due to the size of the company registers, manual annotation is costly, and fine-tuning models with every update in industry classification schemes requires significant data collection. We replicate the manual expert verification by using existing or easily retrievable multimodal resources for industry classification. We present MONETA, the first multimodal industry classification benchmark with text (Website, Wikipedia, Wikidata) and geospatial sources (OpenStreetMap and satellite imagery). Our dataset enlists 1,000 businesses in Europe with 20 economic activity labels according to EU guidelines (NACE). Our training-free baseline reaches 62.10% and 74.10% with open and closed-source Multimodal Large Language Models (MLLM). We observe an increase of up to 22.80% with the combination of multi-turn design, context enrichment, and classification explanations. We will release our dataset and the enhanced guidelines.
title MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems
topic Artificial Intelligence
url https://arxiv.org/abs/2604.07956