PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Kui, Chen, Xinrong, Lv, Wenyu, Liao, Jincheng, Wang, Guanzhong, Liu, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
by: Ni, Feng, et al.
Published: (2025)
by: Ni, Feng, et al.
Published: (2025)
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
by: Lv, Wenyu, et al.
Published: (2024)
by: Lv, Wenyu, et al.
Published: (2024)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
by: Feng, Xiang, et al.
Published: (2026)
by: Feng, Xiang, et al.
Published: (2026)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
by: Hong, Jixiang, et al.
Published: (2025)
by: Hong, Jixiang, et al.
Published: (2025)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
by: Ma, Yubo, et al.
Published: (2024)
by: Ma, Yubo, et al.
Published: (2024)
Docopilot: Improving Multimodal Models for Document-Level Understanding
by: Duan, Yuchen, et al.
Published: (2025)
by: Duan, Yuchen, et al.
Published: (2025)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
by: Tanaka, Ryota, et al.
Published: (2024)
by: Tanaka, Ryota, et al.
Published: (2024)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
by: Zhu, Dawei, et al.
Published: (2025)
by: Zhu, Dawei, et al.
Published: (2025)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
by: Heakl, Ahmed, et al.
Published: (2026)
by: Heakl, Ahmed, et al.
Published: (2026)
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
by: Al-Homoud, Haneen, et al.
Published: (2025)
by: Al-Homoud, Haneen, et al.
Published: (2025)
DocReward: A Document Reward Model for Structuring and Stylizing
by: Liu, Junpeng, et al.
Published: (2025)
by: Liu, Junpeng, et al.
Published: (2025)
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
by: Chau, Phan Phuong Mai, et al.
Published: (2024)
by: Chau, Phan Phuong Mai, et al.
Published: (2024)
DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
by: Liao, Wenhui, et al.
Published: (2024)
by: Liao, Wenhui, et al.
Published: (2024)
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding
by: Le, Binh M., et al.
Published: (2025)
by: Le, Binh M., et al.
Published: (2025)
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
by: Hannan, Tanveer, et al.
Published: (2025)
by: Hannan, Tanveer, et al.
Published: (2025)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
by: Zhu, Wang, et al.
Published: (2023)
by: Zhu, Wang, et al.
Published: (2023)
PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction
by: Sun, Ting, et al.
Published: (2025)
by: Sun, Ting, et al.
Published: (2025)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
by: Gao, Sensen, et al.
Published: (2025)
by: Gao, Sensen, et al.
Published: (2025)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
by: Wang, Wenjie, et al.
Published: (2026)
by: Wang, Wenjie, et al.
Published: (2026)
Effective Training Data Synthesis for Improving MLLM Chart Understanding
by: Yang, Yuwei, et al.
Published: (2025)
by: Yang, Yuwei, et al.
Published: (2025)
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
by: Yang, Minglai, et al.
Published: (2026)
by: Yang, Minglai, et al.
Published: (2026)
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
by: Wang, An-Lan, et al.
Published: (2025)
by: Wang, An-Lan, et al.
Published: (2025)
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
by: Chen, Xiuwei, et al.
Published: (2025)
by: Chen, Xiuwei, et al.
Published: (2025)
DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting
by: Islam, Md Mofijul, et al.
Published: (2026)
by: Islam, Md Mofijul, et al.
Published: (2026)
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
by: Niu, Yuwei, et al.
Published: (2025)
by: Niu, Yuwei, et al.
Published: (2025)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
by: Wu, Te-Lin, et al.
Published: (2021)
by: Wu, Te-Lin, et al.
Published: (2021)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
by: Ding, Chuanghao, et al.
Published: (2024)
by: Ding, Chuanghao, et al.
Published: (2024)
DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
by: Feng, Hao, et al.
Published: (2023)
by: Feng, Hao, et al.
Published: (2023)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents
by: Qian, Kun, et al.
Published: (2025)
by: Qian, Kun, et al.
Published: (2025)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
by: Zhao, Tiancheng, et al.
Published: (2024)
by: Zhao, Tiancheng, et al.
Published: (2024)
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
by: Chen, Guo, et al.
Published: (2026)
by: Chen, Guo, et al.
Published: (2026)
Efficient Document Parsing via Parallel Token Prediction
by: Li, Lei, et al.
Published: (2026)
by: Li, Lei, et al.
Published: (2026)
Multimodal Information Fusion for Chart Understanding: A Survey of MLLMs -- Evolution, Limitations, and Cognitive Enhancement
by: Yi, Zhihang, et al.
Published: (2026)
by: Yi, Zhihang, et al.
Published: (2026)
Similar Items
-
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
by: Ni, Feng, et al.
Published: (2025) -
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
by: Lv, Wenyu, et al.
Published: (2024) -
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
by: Feng, Xiang, et al.
Published: (2026) -
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
by: Kim, Sungnyun, et al.
Published: (2024) -
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
by: Hong, Jixiang, et al.
Published: (2025)