WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, An-Lan, Tang, Jingqun, Lei, Liao, Feng, Hao, Liu, Qi, Fei, Xiang, Lu, Jinghui, Wang, Han, Liu, Weiwei, Liu, Hao, Liu, Yuliang, Bai, Xiang, Huang, Can |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
by: Feng, Hao, et al.
Published: (2023)
by: Feng, Hao, et al.
Published: (2023)
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
by: Yan, Hao, et al.
Published: (2026)
by: Yan, Hao, et al.
Published: (2026)
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
by: Feng, Hao, et al.
Published: (2025)
by: Feng, Hao, et al.
Published: (2025)
Advancing Sequential Numerical Prediction in Autoregressive Models
by: Fei, Xiang, et al.
Published: (2025)
by: Fei, Xiang, et al.
Published: (2025)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
by: Lu, Jinghui, et al.
Published: (2024)
by: Lu, Jinghui, et al.
Published: (2024)
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
by: Zhou, Changda, et al.
Published: (2026)
by: Zhou, Changda, et al.
Published: (2026)
TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering
by: Zhu, Hanshen, et al.
Published: (2026)
by: Zhu, Hanshen, et al.
Published: (2026)
TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
by: Zhao, Weichao, et al.
Published: (2024)
by: Zhao, Weichao, et al.
Published: (2024)
MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
by: Shan, Bin, et al.
Published: (2024)
by: Shan, Bin, et al.
Published: (2024)
Benchmarking Table Comprehension In The Wild
by: Pan, Yikang, et al.
Published: (2024)
by: Pan, Yikang, et al.
Published: (2024)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
by: Ding, Chuanghao, et al.
Published: (2024)
by: Ding, Chuanghao, et al.
Published: (2024)
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
by: Liu, Yuliang, et al.
Published: (2024)
by: Liu, Yuliang, et al.
Published: (2024)
Visual Text Generation in the Wild
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
by: Peng, Hao, et al.
Published: (2026)
by: Peng, Hao, et al.
Published: (2026)
MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
by: Tang, Jingqun, et al.
Published: (2024)
by: Tang, Jingqun, et al.
Published: (2024)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
by: Huang, Kui, et al.
Published: (2025)
by: Huang, Kui, et al.
Published: (2025)
Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
by: Feng, Hao, et al.
Published: (2026)
by: Feng, Hao, et al.
Published: (2026)
WildFusion: Multimodal Implicit 3D Reconstructions in the Wild
by: Liu, Yanbaihui, et al.
Published: (2024)
by: Liu, Yanbaihui, et al.
Published: (2024)
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
by: Tang, Jingqun, et al.
Published: (2024)
by: Tang, Jingqun, et al.
Published: (2024)
LLMs for Relational Reasoning: How Far are We?
by: Li, Zhiming, et al.
Published: (2024)
by: Li, Zhiming, et al.
Published: (2024)
WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
by: Liu, Tengxiao, et al.
Published: (2026)
by: Liu, Tengxiao, et al.
Published: (2026)
How Far are LLMs from Real Search? A Comprehensive Study on Efficiency, Completeness, and Inherent Capabilities
by: Lin, Minhua, et al.
Published: (2025)
by: Lin, Minhua, et al.
Published: (2025)
Partial Scene Text Retrieval
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Model Editing for LLMs4Code: How Far are We?
by: Li, Xiaopeng, et al.
Published: (2024)
by: Li, Xiaopeng, et al.
Published: (2024)
Wild2Avatar: Rendering Humans Behind Occlusions
by: Xiang, Tiange, et al.
Published: (2023)
by: Xiang, Tiange, et al.
Published: (2023)
MedHorizon: Towards Long-context Medical Video Understanding in the Wild
by: Du, Bodong, et al.
Published: (2026)
by: Du, Bodong, et al.
Published: (2026)
Edicho: Consistent Image Editing in the Wild
by: Bai, Qingyan, et al.
Published: (2024)
by: Bai, Qingyan, et al.
Published: (2024)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
by: Feng, Xiang, et al.
Published: (2026)
by: Feng, Xiang, et al.
Published: (2026)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
by: Deng, Chao, et al.
Published: (2024)
by: Deng, Chao, et al.
Published: (2024)
WildChat: 1M ChatGPT Interaction Logs in the Wild
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
by: van der Maden, Willem, et al.
Published: (2026)
by: van der Maden, Willem, et al.
Published: (2026)
How Far Are We From AGI: Are LLMs All We Need?
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
WildLMa: Long Horizon Loco-Manipulation in the Wild
by: Qiu, Ri-Zhao, et al.
Published: (2024)
by: Qiu, Ri-Zhao, et al.
Published: (2024)
Static Application Security Testing (SAST) Tools for Smart Contracts: How Far Are We?
by: Li, Kaixuan, et al.
Published: (2024)
by: Li, Kaixuan, et al.
Published: (2024)
Deep Learning-Based Identification of Inconsistent Method Names: How Far Are We?
by: Wang, Taiming, et al.
Published: (2025)
by: Wang, Taiming, et al.
Published: (2025)
DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies
by: Tao, Tony, et al.
Published: (2025)
by: Tao, Tony, et al.
Published: (2025)
How Far Are We from Genuinely Useful Deep Research Agents?
by: Zhang, Dingling, et al.
Published: (2025)
by: Zhang, Dingling, et al.
Published: (2025)
Adaptive Contextual Embedding for Robust Far-View Borehole Detection
by: Liu, Xuesong, et al.
Published: (2025)
by: Liu, Xuesong, et al.
Published: (2025)
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
by: Deng, Yuntian, et al.
Published: (2024)
by: Deng, Yuntian, et al.
Published: (2024)
Similar Items
-
DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
by: Feng, Hao, et al.
Published: (2023) -
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
by: Yan, Hao, et al.
Published: (2026) -
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
by: Yu, Wenwen, et al.
Published: (2025) -
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
by: Feng, Hao, et al.
Published: (2025) -
Advancing Sequential Numerical Prediction in Autoregressive Models
by: Fei, Xiang, et al.
Published: (2025)