From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yandi, Zhan, Libin, Huang, Ziwei, Luo, Tiancheng, Jiang, Yuxuan, Dong, Wang, Gan, Leilei, Chen, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enterprise Receipt Corpus: 13 Annotated Real-World Receipts for Document AI Pipeline Evaluation
by: Karakavalasa, Naveen
Published: (2026)
by: Karakavalasa, Naveen
Published: (2026)
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
by: Huang, Ziwei, et al.
Published: (2024)
by: Huang, Ziwei, et al.
Published: (2024)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
by: Cai, Yuxuan, et al.
Published: (2025)
by: Cai, Yuxuan, et al.
Published: (2025)
UAVBench and UAVIT-1M: Benchmarking and Enhancing MLLMs for Low-Altitude UAV Vision-Language Understanding
by: Zhan, Yang, et al.
Published: (2026)
by: Zhan, Yang, et al.
Published: (2026)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
by: Li, Chentao, et al.
Published: (2026)
by: Li, Chentao, et al.
Published: (2026)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
by: Cheng, Ziming, et al.
Published: (2025)
by: Cheng, Ziming, et al.
Published: (2025)
ReceiptSense: Beyond Traditional OCR -- A Dataset for Receipt Understanding
by: Abdallah, Abdelrahman, et al.
Published: (2024)
by: Abdallah, Abdelrahman, et al.
Published: (2024)
STAR: A Benchmark for Situated Reasoning in Real-World Videos
by: Wu, Bo, et al.
Published: (2024)
by: Wu, Bo, et al.
Published: (2024)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
by: Xiao, Wenyi, et al.
Published: (2025)
by: Xiao, Wenyi, et al.
Published: (2025)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
by: Hao, Yunzhuo, et al.
Published: (2025)
by: Hao, Yunzhuo, et al.
Published: (2025)
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
by: Zhu, Xiaorong, et al.
Published: (2025)
by: Zhu, Xiaorong, et al.
Published: (2025)
From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company
by: Yu, Zhengxu, et al.
Published: (2026)
by: Yu, Zhengxu, et al.
Published: (2026)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
by: Feng, Xiang, et al.
Published: (2026)
by: Feng, Xiang, et al.
Published: (2026)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025)
by: Fang, I-Sheng, et al.
Published: (2025)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
by: Jiang, Shixin, et al.
Published: (2024)
by: Jiang, Shixin, et al.
Published: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
by: Yilmaz, Nilay, et al.
Published: (2025)
by: Yilmaz, Nilay, et al.
Published: (2025)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
by: Yuan, Jiakang, et al.
Published: (2025)
by: Yuan, Jiakang, et al.
Published: (2025)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
by: Huang, Jincai, et al.
Published: (2026)
by: Huang, Jincai, et al.
Published: (2026)
Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning
by: Ren, Houxing, et al.
Published: (2026)
by: Ren, Houxing, et al.
Published: (2026)
Not Search, But Scan: Benchmarking MLLMs on Scan-Oriented Academic Paper Reasoning
by: Li, Rongjin, et al.
Published: (2026)
by: Li, Rongjin, et al.
Published: (2026)
Refining the Understanding of Autoimmune Liver Disease Mortality
by: Lingyun Zhu, et al.
Published: (2025)
by: Lingyun Zhu, et al.
Published: (2025)
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
by: Lu, Shuo, et al.
Published: (2026)
by: Lu, Shuo, et al.
Published: (2026)
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation
by: Huang, Ziwei, et al.
Published: (2025)
by: Huang, Ziwei, et al.
Published: (2025)
Beyond Unimodal Shortcuts: MLLMs as Cross-Modal Reasoners for Grounded Named Entity Recognition
by: Ma, Jinlong, et al.
Published: (2026)
by: Ma, Jinlong, et al.
Published: (2026)
Showing the Receipts: Understanding the Modern Ransomware Ecosystem
by: Cable, Jack, et al.
Published: (2024)
by: Cable, Jack, et al.
Published: (2024)
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
by: Ye, Hengwei, et al.
Published: (2026)
by: Ye, Hengwei, et al.
Published: (2026)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
CSR-Bench: A Benchmark for Evaluating the Cross-modal Safety and Reliability of MLLMs
by: Liu, Yuxuan, et al.
Published: (2026)
by: Liu, Yuxuan, et al.
Published: (2026)
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training
by: Jiang, Zhouqiang, et al.
Published: (2024)
by: Jiang, Zhouqiang, et al.
Published: (2024)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
by: Xu, Pengju, et al.
Published: (2025)
by: Xu, Pengju, et al.
Published: (2025)
VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
by: Xiao, Wenyi, et al.
Published: (2026)
by: Xiao, Wenyi, et al.
Published: (2026)
CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery Planning
by: Meng, Yang, et al.
Published: (2025)
by: Meng, Yang, et al.
Published: (2025)
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
by: Zhang, Yuanhan, et al.
Published: (2025)
by: Zhang, Yuanhan, et al.
Published: (2025)
Geometric inequalities for convex spacelike hypersurface in de Sitter space
by: Dong, Yandi, et al.
Published: (2025)
by: Dong, Yandi, et al.
Published: (2025)
OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
by: Shen, Jiyuan, et al.
Published: (2026)
by: Shen, Jiyuan, et al.
Published: (2026)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
Similar Items
-
Enterprise Receipt Corpus: 13 Annotated Real-World Receipts for Document AI Pipeline Evaluation
by: Karakavalasa, Naveen
Published: (2026) -
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
by: Huang, Ziwei, et al.
Published: (2024) -
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
by: Cai, Yuxuan, et al.
Published: (2025) -
UAVBench and UAVIT-1M: Benchmarking and Enhancing MLLMs for Low-Altitude UAV Vision-Language Understanding
by: Zhan, Yang, et al.
Published: (2026) -
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
by: Li, Chentao, et al.
Published: (2026)