Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Denghao, Liu, Qing, Chen, Zulong, Xu, Chuanfei, Xu, Jia, Yang, Zhibo, Shao, Wei, Li, Zhao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917495460331520
author Ma, Denghao
Liu, Qing
Chen, Zulong
Xu, Chuanfei
Xu, Jia
Yang, Zhibo
Shao, Wei
Li, Zhao
author_facet Ma, Denghao
Liu, Qing
Chen, Zulong
Xu, Chuanfei
Xu, Jia
Yang, Zhibo
Shao, Wei
Li, Zhao
contents Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -- single domain settings with flat label structures -- that bear little resemblance to the hierarchical, multi-modal, and cross-domain nature of real-world business documents. This gap not only misrepresents practical complexity but also stifles progress toward industrially viable document intelligence. To bridge this gap, we construct the first Multi-level, Multi-domain, Multi-modal document classification Benchmark (MMM-Bench). MMM-Bench includes (1) a deeply hierarchical taxonomy spanning five levels that capture the authentic organizational logic of business documentation; and (2) 5,990 real-world multi-modal documents meticulously curated from 12 commercial domains in Alibaba. Each document is manually annotated with a complete hierarchical path by domain experts. We establish comprehensive baselines on MMM-Bench, which consists of open-weight models and API-based models. Through systematic experiments, we identify four fundamental challenges within MMM-Bench and propose corresponding insights. To provide a solid foundation for advancing research in multi-level, multi-domain document classification, we release all of the data and the evaluation toolkit of MMM-Bench at https://github.com/MMMDC-Bench/MMMDC-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10550
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy
Ma, Denghao
Liu, Qing
Chen, Zulong
Xu, Chuanfei
Xu, Jia
Yang, Zhibo
Shao, Wei
Li, Zhao
Computation and Language
Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -- single domain settings with flat label structures -- that bear little resemblance to the hierarchical, multi-modal, and cross-domain nature of real-world business documents. This gap not only misrepresents practical complexity but also stifles progress toward industrially viable document intelligence. To bridge this gap, we construct the first Multi-level, Multi-domain, Multi-modal document classification Benchmark (MMM-Bench). MMM-Bench includes (1) a deeply hierarchical taxonomy spanning five levels that capture the authentic organizational logic of business documentation; and (2) 5,990 real-world multi-modal documents meticulously curated from 12 commercial domains in Alibaba. Each document is manually annotated with a complete hierarchical path by domain experts. We establish comprehensive baselines on MMM-Bench, which consists of open-weight models and API-based models. Through systematic experiments, we identify four fundamental challenges within MMM-Bench and propose corresponding insights. To provide a solid foundation for advancing research in multi-level, multi-domain document classification, we release all of the data and the evaluation toolkit of MMM-Bench at https://github.com/MMMDC-Bench/MMMDC-Bench.
title Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy
topic Computation and Language
url https://arxiv.org/abs/2605.10550