Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Denghao, Liu, Qing, Chen, Zulong, Xu, Chuanfei, Xu, Jia, Yang, Zhibo, Shao, Wei, Li, Zhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translation
von: Man, Zhibo, et al.
Veröffentlicht: (2025)
von: Man, Zhibo, et al.
Veröffentlicht: (2025)
CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
von: Xu, Zhipeng, et al.
Veröffentlicht: (2026)
von: Xu, Zhipeng, et al.
Veröffentlicht: (2026)
FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
von: Chen, Qiguang, et al.
Veröffentlicht: (2024)
von: Chen, Qiguang, et al.
Veröffentlicht: (2024)
UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
Revisiting Classification Taxonomy for Grammatical Errors
von: Zou, Deqing, et al.
Veröffentlicht: (2025)
von: Zou, Deqing, et al.
Veröffentlicht: (2025)
Multi-modal Stance Detection: New Datasets and Model
von: Liang, Bin, et al.
Veröffentlicht: (2024)
von: Liang, Bin, et al.
Veröffentlicht: (2024)
VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning
von: Jiang, Zhihuan, et al.
Veröffentlicht: (2024)
von: Jiang, Zhihuan, et al.
Veröffentlicht: (2024)
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
CaseGen: A Benchmark for Multi-Stage Legal Case Documents Generation
von: Li, Haitao, et al.
Veröffentlicht: (2025)
von: Li, Haitao, et al.
Veröffentlicht: (2025)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding
von: Jiang, Yue, et al.
Veröffentlicht: (2025)
von: Jiang, Yue, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for Conversational Question Answering in Multi-instructional Documents
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization
von: Ye, Yangfan, et al.
Veröffentlicht: (2024)
von: Ye, Yangfan, et al.
Veröffentlicht: (2024)
CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models
von: Li, Zhong-Zhi, et al.
Veröffentlicht: (2024)
von: Li, Zhong-Zhi, et al.
Veröffentlicht: (2024)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
MultiADE: A Multi-domain Benchmark for Adverse Drug Event Extraction
von: Dai, Xiang, et al.
Veröffentlicht: (2024)
von: Dai, Xiang, et al.
Veröffentlicht: (2024)
Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models
von: Zhao, Shitian, et al.
Veröffentlicht: (2024)
von: Zhao, Shitian, et al.
Veröffentlicht: (2024)
A Survey on Multi-modal Machine Translation: Tasks, Methods and Challenges
von: Shen, Huangjun, et al.
Veröffentlicht: (2024)
von: Shen, Huangjun, et al.
Veröffentlicht: (2024)
Task Selection and Assignment for Multi-modal Multi-task Dialogue Act Classification with Non-stationary Multi-armed Bandits
von: He, Xiangheng, et al.
Veröffentlicht: (2023)
von: He, Xiangheng, et al.
Veröffentlicht: (2023)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
von: Yu, Shi, et al.
Veröffentlicht: (2024)
von: Yu, Shi, et al.
Veröffentlicht: (2024)
Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
von: Ma, Zi-Ao, et al.
Veröffentlicht: (2024)
von: Ma, Zi-Ao, et al.
Veröffentlicht: (2024)
Instruct-Imagen: Image Generation with Multi-modal Instruction
von: Hu, Hexiang, et al.
Veröffentlicht: (2024)
von: Hu, Hexiang, et al.
Veröffentlicht: (2024)
Improving LLM-based Document-level Machine Translation with Multi-Knowledge Fusion
von: Liu, Bin, et al.
Veröffentlicht: (2025)
von: Liu, Bin, et al.
Veröffentlicht: (2025)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
FaiMA: Feature-aware In-context Learning for Multi-domain Aspect-based Sentiment Analysis
von: Yang, Songhua, et al.
Veröffentlicht: (2024)
von: Yang, Songhua, et al.
Veröffentlicht: (2024)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
von: He, Yun, et al.
Veröffentlicht: (2024)
von: He, Yun, et al.
Veröffentlicht: (2024)
Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
von: Madaan, Divyam, et al.
Veröffentlicht: (2025)
von: Madaan, Divyam, et al.
Veröffentlicht: (2025)
SciEvent: Benchmarking Multi-domain Scientific Event Extraction
von: Dong, Bofu, et al.
Veröffentlicht: (2025)
von: Dong, Bofu, et al.
Veröffentlicht: (2025)
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
MMMOS: Multi-domain Multi-axis Audio Quality Assessment
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
Soft-Prompting with Graph-of-Thought for Multi-modal Representation Learning
von: Yang, Juncheng, et al.
Veröffentlicht: (2024)
von: Yang, Juncheng, et al.
Veröffentlicht: (2024)
MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection
von: Lu, Weihai, et al.
Veröffentlicht: (2026)
von: Lu, Weihai, et al.
Veröffentlicht: (2026)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
Research Experiment on Multi-Model Comparison for Chinese Text Classification Tasks
von: Li, JiaCheng
Veröffentlicht: (2024)
von: Li, JiaCheng
Veröffentlicht: (2024)
Ähnliche Einträge
-
DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translation
von: Man, Zhibo, et al.
Veröffentlicht: (2025) -
CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
von: Xu, Zhipeng, et al.
Veröffentlicht: (2026) -
FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
von: Wang, Yan, et al.
Veröffentlicht: (2025) -
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
von: Ji, Yifan, et al.
Veröffentlicht: (2026) -
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
von: Chen, Qiguang, et al.
Veröffentlicht: (2024)