Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Mo, Ye, Ye, Kai, Mao, Xianwei, Shao, Zirui, Huang, Gang, Zhang, Bo, Xing, Hangdi, Chen, Kehan, Zhou, Huan, Yan, Zixu, Bu, Jiajun, Zhou, Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment
by: Ye, Kai, et al.
Published: (2026)
by: Ye, Kai, et al.
Published: (2026)
Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding
by: Shao, Zirui, et al.
Published: (2024)
by: Shao, Zirui, et al.
Published: (2024)
MaS-VQA: A Mask-and-Select Framework for Knowledge-Based Visual Question Answering
by: Mao, Xianwei, et al.
Published: (2026)
by: Mao, Xianwei, et al.
Published: (2026)
WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
by: Shao, Zirui, et al.
Published: (2024)
by: Shao, Zirui, et al.
Published: (2024)
A Simple yet Effective Layout Token in Large Language Models for Document Understanding
by: Zhu, Zhaoqing, et al.
Published: (2025)
by: Zhu, Zhaoqing, et al.
Published: (2025)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
by: Feng, Xiang, et al.
Published: (2026)
by: Feng, Xiang, et al.
Published: (2026)
BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
by: Huang, Tianyuan, et al.
Published: (2025)
by: Huang, Tianyuan, et al.
Published: (2025)
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
by: Yan, Hao, et al.
Published: (2026)
by: Yan, Hao, et al.
Published: (2026)
AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
by: Lee, Yoonsang, et al.
Published: (2024)
by: Lee, Yoonsang, et al.
Published: (2024)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
by: Ding, Chuanghao, et al.
Published: (2024)
by: Ding, Chuanghao, et al.
Published: (2024)
On walk domination: Between different types of walks and $m_3$-path
by: Chen, Hangdi, et al.
Published: (2025)
by: Chen, Hangdi, et al.
Published: (2025)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
by: Ma, Yubo, et al.
Published: (2024)
by: Ma, Yubo, et al.
Published: (2024)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
CoS: Chain-of-Shot Prompting for Long Video Understanding
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
by: Deng, Chao, et al.
Published: (2024)
by: Deng, Chao, et al.
Published: (2024)
Towards Scalable Web Accessibility Audit with MLLMs as Copilots
by: Gu, Ming, et al.
Published: (2025)
by: Gu, Ming, et al.
Published: (2025)
mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
by: Hu, Anwen, et al.
Published: (2024)
by: Hu, Anwen, et al.
Published: (2024)
PathCoT: Chain-of-Thought Prompting for Zero-shot Pathology Visual Reasoning
by: Zhou, Junjie, et al.
Published: (2025)
by: Zhou, Junjie, et al.
Published: (2025)
VisDocSketcher: Towards Scalable Visual Documentation with Agentic Systems
by: Gomes, Luís F., et al.
Published: (2025)
by: Gomes, Luís F., et al.
Published: (2025)
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models
by: Dai, Xuanlang, et al.
Published: (2026)
by: Dai, Xuanlang, et al.
Published: (2026)
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
by: Hu, Anwen, et al.
Published: (2024)
by: Hu, Anwen, et al.
Published: (2024)
FocusedAD: Character-centric Movie Audio Description
by: Ye, Xiaojun, et al.
Published: (2025)
by: Ye, Xiaojun, et al.
Published: (2025)
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
Revisiting, Benchmarking and Understanding Unsupervised Graph Domain Adaptation
by: Liu, Meihan, et al.
Published: (2024)
by: Liu, Meihan, et al.
Published: (2024)
DocCogito: Aligning Layout Cognition and Step-Level Grounded Reasoning for Document Understanding
by: Wu, Yuchuan, et al.
Published: (2026)
by: Wu, Yuchuan, et al.
Published: (2026)
Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
by: Zhou, Bangbang, et al.
Published: (2026)
by: Zhou, Bangbang, et al.
Published: (2026)
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
by: Chen, Ketong, et al.
Published: (2025)
by: Chen, Ketong, et al.
Published: (2025)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
by: Wang, Wenjie, et al.
Published: (2026)
by: Wang, Wenjie, et al.
Published: (2026)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
by: Tanaka, Ryota, et al.
Published: (2024)
by: Tanaka, Ryota, et al.
Published: (2024)
Investigation of CoB/SBA‐15 Catalyst for Efficient Hydrogen Generation via Sodium Borohydride Hydrolysis
by: Hao Yu, et al.
Published: (2025)
by: Hao Yu, et al.
Published: (2025)
Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
by: Zheng, Yuanlei, et al.
Published: (2026)
by: Zheng, Yuanlei, et al.
Published: (2026)
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
by: Xiong, Junyu, et al.
Published: (2025)
by: Xiong, Junyu, et al.
Published: (2025)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
by: Qian, Jiahe, et al.
Published: (2025)
by: Qian, Jiahe, et al.
Published: (2025)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
by: Zhu, Dawei, et al.
Published: (2025)
by: Zhu, Dawei, et al.
Published: (2025)
Decompiling Rust: An Empirical Study of Compiler Optimizations and Reverse Engineering Challenges
by: Zhou, Zixu
Published: (2025)
by: Zhou, Zixu
Published: (2025)
ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
by: Zhang, Shimin, et al.
Published: (2025)
by: Zhang, Shimin, et al.
Published: (2025)
DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
by: Feng, Hao, et al.
Published: (2023)
by: Feng, Hao, et al.
Published: (2023)
Similar Items
-
REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment
by: Ye, Kai, et al.
Published: (2026) -
Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding
by: Shao, Zirui, et al.
Published: (2024) -
MaS-VQA: A Mask-and-Select Framework for Knowledge-Based Visual Question Answering
by: Mao, Xianwei, et al.
Published: (2026) -
WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
by: Shao, Zirui, et al.
Published: (2024) -
A Simple yet Effective Layout Token in Large Language Models for Document Understanding
by: Zhu, Zhaoqing, et al.
Published: (2025)