dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Yumeng, Yang, Guang, Liu, Hao, Wang, Bowen, Zhang, Colin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multimodal OCR: Parse Anything from Documents
por: Zheng, Handong, et al.
Publicado: (2026)
por: Zheng, Handong, et al.
Publicado: (2026)
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
por: Cui, Cheng, et al.
Publicado: (2025)
por: Cui, Cheng, et al.
Publicado: (2025)
Parser-Oriented Structural Refinement for a Stable Layout Interface in Document Parsing
por: Liu, Fuyuan, et al.
Publicado: (2026)
por: Liu, Fuyuan, et al.
Publicado: (2026)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
por: Li, Zhang, et al.
Publicado: (2026)
por: Li, Zhang, et al.
Publicado: (2026)
HSD: Training-Free Acceleration for Document Parsing Vision-Language Model with Hierarchical Speculative Decoding
por: Liao, Wenhui, et al.
Publicado: (2026)
por: Liao, Wenhui, et al.
Publicado: (2026)
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
por: Niu, Junbo, et al.
Publicado: (2025)
por: Niu, Junbo, et al.
Publicado: (2025)
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training
por: Jiang, Zhouqiang, et al.
Publicado: (2024)
por: Jiang, Zhouqiang, et al.
Publicado: (2024)
GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing
por: Ma, Xianzhi, et al.
Publicado: (2025)
por: Ma, Xianzhi, et al.
Publicado: (2025)
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
por: Nath, Oikantik, et al.
Publicado: (2025)
por: Nath, Oikantik, et al.
Publicado: (2025)
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
por: Wang, Baode, et al.
Publicado: (2025)
por: Wang, Baode, et al.
Publicado: (2025)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
por: Wang, Wenjie, et al.
Publicado: (2026)
por: Wang, Wenjie, et al.
Publicado: (2026)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
por: Luo, Chuwei, et al.
Publicado: (2024)
por: Luo, Chuwei, et al.
Publicado: (2024)
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
por: Sun, Fan-Yun, et al.
Publicado: (2024)
por: Sun, Fan-Yun, et al.
Publicado: (2024)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
por: Li, Zhang, et al.
Publicado: (2025)
por: Li, Zhang, et al.
Publicado: (2025)
ParseBench: A Document Parsing Benchmark for AI Agents
por: Zhang, Boyang, et al.
Publicado: (2026)
por: Zhang, Boyang, et al.
Publicado: (2026)
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
por: Feng, Hao, et al.
Publicado: (2025)
por: Feng, Hao, et al.
Publicado: (2025)
Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
por: Feng, Hao, et al.
Publicado: (2026)
por: Feng, Hao, et al.
Publicado: (2026)
RoDLA: Benchmarking the Robustness of Document Layout Analysis Models
por: Chen, Yufan, et al.
Publicado: (2024)
por: Chen, Yufan, et al.
Publicado: (2024)
PVLM: Parsing-Aware Vision Language Model with Dynamic Contrastive Learning for Zero-Shot Deepfake Attribution
por: Zhang, Yaning, et al.
Publicado: (2025)
por: Zhang, Yaning, et al.
Publicado: (2025)
SciPostLayout: A Dataset for Layout Analysis and Layout Generation of Scientific Posters
por: Tanaka, Shohei, et al.
Publicado: (2024)
por: Tanaka, Shohei, et al.
Publicado: (2024)
Adversarial Supervision Makes Layout-to-Image Diffusion Models Thrive
por: Li, Yumeng, et al.
Publicado: (2024)
por: Li, Yumeng, et al.
Publicado: (2024)
A Simple yet Effective Layout Token in Large Language Models for Document Understanding
por: Zhu, Zhaoqing, et al.
Publicado: (2025)
por: Zhu, Zhaoqing, et al.
Publicado: (2025)
Layout-Independent License Plate Recognition via Integrated Vision and Language Models
por: Shabaninia, Elham, et al.
Publicado: (2025)
por: Shabaninia, Elham, et al.
Publicado: (2025)
LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
por: Wu, Yuxuan, et al.
Publicado: (2025)
por: Wu, Yuxuan, et al.
Publicado: (2025)
SpotActor: Training-Free Layout-Controlled Consistent Image Generation
por: Wang, Jiahao, et al.
Publicado: (2024)
por: Wang, Jiahao, et al.
Publicado: (2024)
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
por: Malik, Haq Nawaz, et al.
Publicado: (2026)
por: Malik, Haq Nawaz, et al.
Publicado: (2026)
EvoVLA: Self-Evolving Vision-Language-Action Model
por: Liu, Zeting, et al.
Publicado: (2025)
por: Liu, Zeting, et al.
Publicado: (2025)
HOG-Layout: Hierarchical 3D Scene Generation, Optimization and Editing via Vision-Language Models
por: Jiang, Haiyan, et al.
Publicado: (2026)
por: Jiang, Haiyan, et al.
Publicado: (2026)
HybriDLA: Hybrid Generation for Document Layout Analysis
por: Chen, Yufan, et al.
Publicado: (2025)
por: Chen, Yufan, et al.
Publicado: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
por: Kang, Hengrui, et al.
Publicado: (2025)
por: Kang, Hengrui, et al.
Publicado: (2025)
PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models
por: He, Runze, et al.
Publicado: (2025)
por: He, Runze, et al.
Publicado: (2025)
LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
por: Heo, Inbum, et al.
Publicado: (2025)
por: Heo, Inbum, et al.
Publicado: (2025)
SFDLA: Source-Free Document Layout Analysis
por: Tewes, Sebastian, et al.
Publicado: (2025)
por: Tewes, Sebastian, et al.
Publicado: (2025)
DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
por: Zhang, Qintong, et al.
Publicado: (2025)
por: Zhang, Qintong, et al.
Publicado: (2025)
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
por: Chen, Weixing, et al.
Publicado: (2025)
por: Chen, Weixing, et al.
Publicado: (2025)
PARL: Position-Aware Relation Learning Network for Document Layout Analysis
por: Liu, Fuyuan, et al.
Publicado: (2026)
por: Liu, Fuyuan, et al.
Publicado: (2026)
600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script
por: Malik, Haq Nawaz
Publicado: (2026)
por: Malik, Haq Nawaz
Publicado: (2026)
METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
por: Liu, Yuchen, et al.
Publicado: (2025)
por: Liu, Yuchen, et al.
Publicado: (2025)
LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
por: Lv, Chengtao, et al.
Publicado: (2025)
por: Lv, Chengtao, et al.
Publicado: (2025)
Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation
por: Wang, Xintong, et al.
Publicado: (2025)
por: Wang, Xintong, et al.
Publicado: (2025)
Ejemplares similares
-
Multimodal OCR: Parse Anything from Documents
por: Zheng, Handong, et al.
Publicado: (2026) -
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
por: Cui, Cheng, et al.
Publicado: (2025) -
Parser-Oriented Structural Refinement for a Stable Layout Interface in Document Parsing
por: Liu, Fuyuan, et al.
Publicado: (2026) -
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
por: Li, Zhang, et al.
Publicado: (2026) -
HSD: Training-Free Acceleration for Document Parsing Vision-Language Model with Hierarchical Speculative Decoding
por: Liao, Wenhui, et al.
Publicado: (2026)