Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Hao, Shi, Wei, Zhang, Ke, Fei, Xiang, Liao, Lei, Yang, Dingkang, Du, Yongkun, Wu, Xuecheng, Tang, Jingqun, Liu, Yang, Chen, Hong, Huang, Can |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
by: Feng, Hao, et al.
Published: (2025)
by: Feng, Hao, et al.
Published: (2025)
TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering
by: Zhu, Hanshen, et al.
Published: (2026)
by: Zhu, Hanshen, et al.
Published: (2026)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
by: Xue, Wei, et al.
Published: (2026)
by: Xue, Wei, et al.
Published: (2026)
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
by: Liu, Keliang, et al.
Published: (2025)
by: Liu, Keliang, et al.
Published: (2025)
Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
by: Han, Minghao, et al.
Published: (2025)
by: Han, Minghao, et al.
Published: (2025)
MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
by: Shan, Bin, et al.
Published: (2024)
by: Shan, Bin, et al.
Published: (2024)
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
by: Wang, An-Lan, et al.
Published: (2025)
by: Wang, An-Lan, et al.
Published: (2025)
AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
by: Feng, Hao, et al.
Published: (2023)
by: Feng, Hao, et al.
Published: (2023)
Therapeutic strategies for aberrant splicing in cancer and genetic disorders
by: Wenhua Shi, et al.
Published: (2024)
by: Wenhua Shi, et al.
Published: (2024)
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
by: Yu, Haiyang, et al.
Published: (2025)
by: Yu, Haiyang, et al.
Published: (2025)
ParGo: Bridging Vision-Language with Partial and Global Views
by: Wang, An-Lan, et al.
Published: (2024)
by: Wang, An-Lan, et al.
Published: (2024)
Advancing Sequential Numerical Prediction in Autoregressive Models
by: Fei, Xiang, et al.
Published: (2025)
by: Fei, Xiang, et al.
Published: (2025)
Skip and Skip: Segmenting Medical Images with Prompts
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Parse Trees Guided LLM Prompt Compression
by: Mao, Wenhao, et al.
Published: (2024)
by: Mao, Wenhao, et al.
Published: (2024)
Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
by: Liu, Maoqi, et al.
Published: (2025)
by: Liu, Maoqi, et al.
Published: (2025)
A General Anchor-Based Framework for Scalable Fair Clustering
by: Wei, Shengfei, et al.
Published: (2025)
by: Wei, Shengfei, et al.
Published: (2025)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
by: Nie, Yuxiang, et al.
Published: (2025)
by: Nie, Yuxiang, et al.
Published: (2025)
Dolphin v1.0 Technical Report
by: Weng, Taohan, et al.
Published: (2025)
by: Weng, Taohan, et al.
Published: (2025)
Multimodal OCR: Parse Anything from Documents
by: Zheng, Handong, et al.
Published: (2026)
by: Zheng, Handong, et al.
Published: (2026)
DARTs: A Dual-Path Robust Framework for Anomaly Detection in High-Dimensional Multivariate Time Series
by: Liu, Xuechun, et al.
Published: (2025)
by: Liu, Xuechun, et al.
Published: (2025)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
by: Lu, Jinghui, et al.
Published: (2024)
by: Lu, Jinghui, et al.
Published: (2024)
Diffusion Probe: Generated Image Result Prediction Using CNN Probes
by: Cui, Benlei, et al.
Published: (2026)
by: Cui, Benlei, et al.
Published: (2026)
V2X-RECT: An Efficient V2X Trajectory Prediction Framework via Redundant Interaction Filtering and Tracking Error Correction
by: Kong, Xiangyan, et al.
Published: (2025)
by: Kong, Xiangyan, et al.
Published: (2025)
Dolphin: A Programmable Framework for Scalable Neurosymbolic Learning
by: Naik, Aaditya, et al.
Published: (2024)
by: Naik, Aaditya, et al.
Published: (2024)
TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
by: Zhao, Weichao, et al.
Published: (2024)
by: Zhao, Weichao, et al.
Published: (2024)
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
by: Zhou, Bangbang, et al.
Published: (2026)
by: Zhou, Bangbang, et al.
Published: (2026)
BotzoneBench: Scalable LLM Evaluation via Graded AI Anchors
by: Li, Lingfeng, et al.
Published: (2026)
by: Li, Lingfeng, et al.
Published: (2026)
Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather Removal
by: Liao, Rongxin, et al.
Published: (2025)
by: Liao, Rongxin, et al.
Published: (2025)
MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
by: Xu, Bangrui, et al.
Published: (2026)
by: Xu, Bangrui, et al.
Published: (2026)
dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
by: Li, Yumeng, et al.
Published: (2025)
by: Li, Yumeng, et al.
Published: (2025)
MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
by: Zhang, Jiarui, et al.
Published: (2025)
by: Zhang, Jiarui, et al.
Published: (2025)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
by: Tang, Zihan, et al.
Published: (2026)
by: Tang, Zihan, et al.
Published: (2026)
UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
Towards Comprehensive Interactive Change Understanding in Remote Sensing: A Large-scale Dataset and Dual-granularity Enhanced VLM
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
by: Tang, Jingqun, et al.
Published: (2024)
by: Tang, Jingqun, et al.
Published: (2024)
Dolphin-CN-Dialect: Where Chinese Dialects Matter
by: Meng, Yangyang, et al.
Published: (2026)
by: Meng, Yangyang, et al.
Published: (2026)
Co-PLNet: A Collaborative Point-Line Network for Prompt-Guided Wireframe Parsing
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
FED-Bench: A Cross-Granular Benchmark for Disentangled Evaluation of Facial Expression Editing
by: Xue, Fengjian, et al.
Published: (2026)
by: Xue, Fengjian, et al.
Published: (2026)
Similar Items
-
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
by: Feng, Hao, et al.
Published: (2025) -
TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering
by: Zhu, Hanshen, et al.
Published: (2026) -
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
by: Xue, Wei, et al.
Published: (2026) -
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
by: Du, Yongkun, et al.
Published: (2025) -
Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
by: Liu, Keliang, et al.
Published: (2025)