Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Hao, Wei, Shu, Fei, Xiang, Shi, Wei, Han, Yingdong, Liao, Lei, Lu, Jinghui, Wu, Binghong, Liu, Qi, Lin, Chunhui, Tang, Jingqun, Liu, Hao, Huang, Can |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
by: Feng, Hao, et al.
Published: (2026)
by: Feng, Hao, et al.
Published: (2026)
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
by: Wang, An-Lan, et al.
Published: (2025)
by: Wang, An-Lan, et al.
Published: (2025)
TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
by: Zhao, Weichao, et al.
Published: (2024)
by: Zhao, Weichao, et al.
Published: (2024)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
by: Lu, Jinghui, et al.
Published: (2024)
by: Lu, Jinghui, et al.
Published: (2024)
Harmonizing Visual Text Comprehension and Generation
by: Zhao, Zhen, et al.
Published: (2024)
by: Zhao, Zhen, et al.
Published: (2024)
Advancing Sequential Numerical Prediction in Autoregressive Models
by: Fei, Xiang, et al.
Published: (2025)
by: Fei, Xiang, et al.
Published: (2025)
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
by: Tang, Jingqun, et al.
Published: (2024)
by: Tang, Jingqun, et al.
Published: (2024)
DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
by: Feng, Hao, et al.
Published: (2023)
by: Feng, Hao, et al.
Published: (2023)
Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer
by: Zhao, Zhen, et al.
Published: (2023)
by: Zhao, Zhen, et al.
Published: (2023)
MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
by: Shan, Bin, et al.
Published: (2024)
by: Shan, Bin, et al.
Published: (2024)
MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
by: Tang, Jingqun, et al.
Published: (2024)
by: Tang, Jingqun, et al.
Published: (2024)
ParGo: Bridging Vision-Language with Partial and Global Views
by: Wang, An-Lan, et al.
Published: (2024)
by: Wang, An-Lan, et al.
Published: (2024)
Post-Completion Learning for Language Models
by: Fei, Xiang, et al.
Published: (2025)
by: Fei, Xiang, et al.
Published: (2025)
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
by: Fu, Ling, et al.
Published: (2024)
by: Fu, Ling, et al.
Published: (2024)
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
by: Yu, Haiyang, et al.
Published: (2025)
by: Yu, Haiyang, et al.
Published: (2025)
Therapeutic strategies for aberrant splicing in cancer and genetic disorders
by: Wenhua Shi, et al.
Published: (2024)
by: Wenhua Shi, et al.
Published: (2024)
Adaptive Few-shot Prompting for Machine Translation with Pre-trained Language Models
by: Tang, Lei, et al.
Published: (2025)
by: Tang, Lei, et al.
Published: (2025)
AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering
by: Zhu, Hanshen, et al.
Published: (2026)
by: Zhu, Hanshen, et al.
Published: (2026)
Candidate Pseudolabel Learning: Enhancing Vision-Language Models by Prompt Tuning with Unlabeled Data
by: Zhang, Jiahan, et al.
Published: (2024)
by: Zhang, Jiahan, et al.
Published: (2024)
Multimodal OCR: Parse Anything from Documents
by: Zheng, Handong, et al.
Published: (2026)
by: Zheng, Handong, et al.
Published: (2026)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
by: Lu, Jinghui, et al.
Published: (2025)
by: Lu, Jinghui, et al.
Published: (2025)
MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
by: Jia, Weitao, et al.
Published: (2025)
by: Jia, Weitao, et al.
Published: (2025)
Vision as LoRA
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
by: Li, Yumeng, et al.
Published: (2025)
by: Li, Yumeng, et al.
Published: (2025)
Flexible Multi-Target Angular Emulation for Over-the-Air Testing of Large-Scale ISAC Base Stations: Principle and Experimental Verification
by: Li, Chunhui, et al.
Published: (2026)
by: Li, Chunhui, et al.
Published: (2026)
Against Multifaceted Graph Heterogeneity via Asymmetric Federated Prompt Learning
by: Guo, Zhuoning, et al.
Published: (2024)
by: Guo, Zhuoning, et al.
Published: (2024)
Co-PLNet: A Collaborative Point-Line Network for Prompt-Guided Wireframe Parsing
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
Test-Time Multimodal Backdoor Detection by Contrastive Prompting
by: Niu, Yuwei, et al.
Published: (2024)
by: Niu, Yuwei, et al.
Published: (2024)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
by: Song, Jiahe, et al.
Published: (2025)
by: Song, Jiahe, et al.
Published: (2025)
Efficient Document Parsing via Parallel Token Prediction
by: Li, Lei, et al.
Published: (2026)
by: Li, Lei, et al.
Published: (2026)
Optimizing Prompts for Text-to-Image Generation
by: Hao, Yaru, et al.
Published: (2022)
by: Hao, Yaru, et al.
Published: (2022)
Stability Anchors and Risk Amplifiers: Tail Spillovers Across Stablecoin Designs
by: Wu, Wenbin, et al.
Published: (2026)
by: Wu, Wenbin, et al.
Published: (2026)
Optimal Anchor Deployment and Topology Design for Large-Scale AUV Navigation
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
by: Yan, Hao, et al.
Published: (2026)
by: Yan, Hao, et al.
Published: (2026)
Parse Trees Guided LLM Prompt Compression
by: Mao, Wenhao, et al.
Published: (2024)
by: Mao, Wenhao, et al.
Published: (2024)
Revisiting Structured Sentiment Analysis as Latent Dependency Graph Parsing
by: Zhou, Chengjie, et al.
Published: (2024)
by: Zhou, Chengjie, et al.
Published: (2024)
Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
by: Liu, Maoqi, et al.
Published: (2025)
by: Liu, Maoqi, et al.
Published: (2025)
Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
by: Liu, Keliang, et al.
Published: (2025)
by: Liu, Keliang, et al.
Published: (2025)
GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing
by: Ma, Xianzhi, et al.
Published: (2025)
by: Ma, Xianzhi, et al.
Published: (2025)
Similar Items
-
Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
by: Feng, Hao, et al.
Published: (2026) -
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
by: Wang, An-Lan, et al.
Published: (2025) -
TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
by: Zhao, Weichao, et al.
Published: (2024) -
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
by: Lu, Jinghui, et al.
Published: (2024) -
Harmonizing Visual Text Comprehension and Generation
by: Zhao, Zhen, et al.
Published: (2024)