A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Yihao, Luo, Siwen, Dai, Yue, Jiang, Yanbei, Li, Zechuan, Sun, Qiang, Martin, Geoffrey, Liu, Wei, Peng, Yifan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding
by: Ding, Yihao, et al.
Published: (2026)
by: Ding, Yihao, et al.
Published: (2026)
SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends
by: Barboule, Camille, et al.
Published: (2025)
by: Barboule, Camille, et al.
Published: (2025)
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
by: Li, Zechuan, et al.
Published: (2025)
by: Li, Zechuan, et al.
Published: (2025)
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
PROPA: Toward Process-level Optimization in Visual Reasoning via Reinforcement Learning
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
SynJAC: Synthetic-data-driven Joint-granular Adaptation and Calibration for Domain Specific Scanned Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
by: Kuang, Jiayi, et al.
Published: (2024)
by: Kuang, Jiayi, et al.
Published: (2024)
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
Bootstrap Method in Theoretical Physics
by: Zheng, Zechuan
Published: (2023)
by: Zheng, Zechuan
Published: (2023)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
by: Luo, Chuwei, et al.
Published: (2022)
by: Luo, Chuwei, et al.
Published: (2022)
Harnessing Webpage UIs for Text-Rich Visual Understanding
by: Liu, Junpeng, et al.
Published: (2024)
by: Liu, Junpeng, et al.
Published: (2024)
Multi-Intent Spoken Language Understanding: Methods, Trends, and Challenges
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Semantic Communication Meets Heterogeneous Network: Emerging Trends, Opportunities, and Challenges
by: Zheng, Guhan, et al.
Published: (2025)
by: Zheng, Guhan, et al.
Published: (2025)
Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
by: Xu, Zhenhua, et al.
Published: (2025)
by: Xu, Zhenhua, et al.
Published: (2025)
Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey
by: Zhang, Xiantao
Published: (2025)
by: Zhang, Xiantao
Published: (2025)
The Healing Potential of Platelet‐Rich Plasma: Advances in Preparation Methods, Biomedical Applications, and Emerging Challenges
by: Mustafijur Rahman, et al.
Published: (2025)
by: Mustafijur Rahman, et al.
Published: (2025)
Surveying the MLLM Landscape: A Meta-Review of Current Surveys
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector
by: Li, Zechuan, et al.
Published: (2025)
by: Li, Zechuan, et al.
Published: (2025)
LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document Understanding
by: Sourati, Zhivar, et al.
Published: (2025)
by: Sourati, Zhivar, et al.
Published: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
Relation-Rich Visual Document Generator for Visual Information Extraction
by: Jiang, Zi-Han, et al.
Published: (2025)
by: Jiang, Zi-Han, et al.
Published: (2025)
MARCH: Multi-Agent Radiology Clinical Hierarchy for CT Report Generation
by: Lin, Yi, et al.
Published: (2026)
by: Lin, Yi, et al.
Published: (2026)
Local Interpretations for Explainable Natural Language Processing: A Survey
by: Luo, Siwen, et al.
Published: (2021)
by: Luo, Siwen, et al.
Published: (2021)
Emerging Trends, Approaches And Challenges In Engineering Education In The UK
by: Fowler, Stella, et al.
Published: (2023)
by: Fowler, Stella, et al.
Published: (2023)
Emerging Trends and Challenges in Supply Chain Management and Sustainability
by: Isabel Castillo‐Pérez, et al.
Published: (2025)
by: Isabel Castillo‐Pérez, et al.
Published: (2025)
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
by: Chen, Ketong, et al.
Published: (2025)
by: Chen, Ketong, et al.
Published: (2025)
LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
by: Fujitake, Masato
Published: (2024)
by: Fujitake, Masato
Published: (2024)
Budget-Aware Routing for Long Clinical Text
by: Qureshi, Khizar, et al.
Published: (2026)
by: Qureshi, Khizar, et al.
Published: (2026)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
by: Wu, Mingrui, et al.
Published: (2024)
by: Wu, Mingrui, et al.
Published: (2024)
Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects
by: Wu, Chengyan, et al.
Published: (2025)
by: Wu, Chengyan, et al.
Published: (2025)
Survey on AI-Generated Media Detection: From Non-MLLM to MLLM
by: Zou, Yueying, et al.
Published: (2025)
by: Zou, Yueying, et al.
Published: (2025)
A Survey on Game Playing Agents and Large Models: Methods, Applications, and Challenges
by: Xu, Xinrun, et al.
Published: (2024)
by: Xu, Xinrun, et al.
Published: (2024)
Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language Models
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
Bibliometric Analysis of Charcot Arthropathy (1995–2025): Current Status and Emerging Trends
by: Jian Lin Zhou, et al.
Published: (2026)
by: Jian Lin Zhou, et al.
Published: (2026)
Similar Items
-
Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding
by: Ding, Yihao, et al.
Published: (2026) -
SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2025) -
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024) -
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends
by: Barboule, Camille, et al.
Published: (2025) -
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
by: Ding, Yihao, et al.
Published: (2025)