VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Yihao, Han, Soyeon Caren, Li, Yan, Poon, Josiah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
SynJAC: Synthetic-data-driven Joint-granular Adaptation and Calibration for Domain Specific Scanned Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
Location-Aware Pretraining for Medical Difference Visual Question Answering
by: Musinguzi, Denis, et al.
Published: (2026)
by: Musinguzi, Denis, et al.
Published: (2026)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
VERSE: Visual Embedding Reduction and Space Exploration. Clustering-Guided Insights for Training Data Enhancement in Visually-Rich Document Understanding
by: de Rodrigo, Ignacio, et al.
Published: (2026)
by: de Rodrigo, Ignacio, et al.
Published: (2026)
Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich Documents
by: Napolitano, Davide, et al.
Published: (2025)
by: Napolitano, Davide, et al.
Published: (2025)
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
by: Fujitake, Masato
Published: (2024)
by: Fujitake, Masato
Published: (2024)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
by: Mo, Ye, et al.
Published: (2025)
by: Mo, Ye, et al.
Published: (2025)
Automatic Layout Planning for Visually-Rich Documents with Instruction-Following Models
by: Zhu, Wanrong, et al.
Published: (2024)
by: Zhu, Wanrong, et al.
Published: (2024)
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments
by: Nguyen, Van Quang
Published: (2026)
by: Nguyen, Van Quang
Published: (2026)
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
by: Wang, Qiuchen, et al.
Published: (2025)
by: Wang, Qiuchen, et al.
Published: (2025)
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
by: Van Landeghem, Jordy, et al.
Published: (2024)
by: Van Landeghem, Jordy, et al.
Published: (2024)
Internalized Reasoning for Long-Context Visual Document Understanding
by: Veselka, Austin
Published: (2026)
by: Veselka, Austin
Published: (2026)
Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich Documents
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
by: Tien, Dong Nguyen, et al.
Published: (2025)
by: Tien, Dong Nguyen, et al.
Published: (2025)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
by: Zhu, Fengbin, et al.
Published: (2024)
by: Zhu, Fengbin, et al.
Published: (2024)
CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
by: Han, Hongyong, et al.
Published: (2025)
by: Han, Hongyong, et al.
Published: (2025)
MARCH: Multi-Agent Radiology Clinical Hierarchy for CT Report Generation
by: Lin, Yi, et al.
Published: (2026)
by: Lin, Yi, et al.
Published: (2026)
ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales
by: Zhen, Yihao, et al.
Published: (2025)
by: Zhen, Yihao, et al.
Published: (2025)
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
by: Kawasaki, Haruka, et al.
Published: (2026)
by: Kawasaki, Haruka, et al.
Published: (2026)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
by: Tanaka, Ryota, et al.
Published: (2025)
by: Tanaka, Ryota, et al.
Published: (2025)
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
by: Li, Yueying, et al.
Published: (2026)
by: Li, Yueying, et al.
Published: (2026)
Text-guided 3D Human Motion Generation with Keyframe-based Parallel Skip Transformer
by: Geng, Zichen, et al.
Published: (2024)
by: Geng, Zichen, et al.
Published: (2024)
Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence
by: Rau, Anita, et al.
Published: (2025)
by: Rau, Anita, et al.
Published: (2025)
Towards Understanding Visual Grounding in Visual Language Models
by: Pantazopoulos, Georgios, et al.
Published: (2025)
by: Pantazopoulos, Georgios, et al.
Published: (2025)
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
by: Dai, Ming, et al.
Published: (2024)
by: Dai, Ming, et al.
Published: (2024)
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
by: Kwon, Mincheol, et al.
Published: (2026)
by: Kwon, Mincheol, et al.
Published: (2026)
Boximator: Generating Rich and Controllable Motions for Video Synthesis
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
Dual Latent Memory for Visual Multi-agent System
by: Yu, Xinlei, et al.
Published: (2026)
by: Yu, Xinlei, et al.
Published: (2026)
MSG-Chart: Multimodal Scene Graph for ChartQA
by: Dai, Yue, et al.
Published: (2024)
by: Dai, Yue, et al.
Published: (2024)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
DOGR: Towards Versatile Visual Document Grounding and Referring
by: Zhou, Yinan, et al.
Published: (2024)
by: Zhou, Yinan, et al.
Published: (2024)
Similar Items
-
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
by: Wang, Eileen, et al.
Published: (2024) -
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024) -
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
by: Ding, Yihao, et al.
Published: (2024) -
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025) -
SynJAC: Synthetic-data-driven Joint-granular Adaptation and Calibration for Domain Specific Scanned Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2024)