Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Baode, Wu, Biao, Li, Weizhen, Fang, Meng, Huang, Zuming, Huang, Jun, Wang, Haozhe, Liang, Yanjie, Chen, Ling, Chu, Wei, Qi, Yuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
by: Wang, Baode, et al.
Published: (2025)
by: Wang, Baode, et al.
Published: (2025)
Parser-Oriented Structural Refinement for a Stable Layout Interface in Document Parsing
by: Liu, Fuyuan, et al.
Published: (2026)
by: Liu, Fuyuan, et al.
Published: (2026)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
by: Jing, Hongyi, et al.
Published: (2025)
by: Jing, Hongyi, et al.
Published: (2025)
UniParser: Multi-Human Parsing with Unified Correlation Representation Learning
by: Chu, Jiaming, et al.
Published: (2023)
by: Chu, Jiaming, et al.
Published: (2023)
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding
by: Xu, Pengxin, et al.
Published: (2026)
by: Xu, Pengxin, et al.
Published: (2026)
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
CREPE: Coordinate-Aware End-to-End Document Parser
by: Okamoto, Yamato, et al.
Published: (2024)
by: Okamoto, Yamato, et al.
Published: (2024)
dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
by: Li, Yumeng, et al.
Published: (2025)
by: Li, Yumeng, et al.
Published: (2025)
Uni-Parser Technical Report
by: Fang, Xi, et al.
Published: (2025)
by: Fang, Xi, et al.
Published: (2025)
How Do Document Parsers Break? Auditing Structural Vulnerability in Document Intelligence
by: Chen, Yue, et al.
Published: (2026)
by: Chen, Yue, et al.
Published: (2026)
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers
by: Chen, Weidong, et al.
Published: (2026)
by: Chen, Weidong, et al.
Published: (2026)
PARL: Position-Aware Relation Learning Network for Document Layout Analysis
by: Liu, Fuyuan, et al.
Published: (2026)
by: Liu, Fuyuan, et al.
Published: (2026)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
by: Zhang, Qintong, et al.
Published: (2024)
by: Zhang, Qintong, et al.
Published: (2024)
LAPDoc: Layout-Aware Prompting for Documents
by: Lamott, Marcel, et al.
Published: (2024)
by: Lamott, Marcel, et al.
Published: (2024)
Vision-Language Reasoning for Geolocalization: A Reinforcement Learning Approach
by: Wu, Biao, et al.
Published: (2026)
by: Wu, Biao, et al.
Published: (2026)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
PresentAgent: Multimodal Agent for Presentation Video Generation
by: Shi, Jingwei, et al.
Published: (2025)
by: Shi, Jingwei, et al.
Published: (2025)
VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Efficient Document Parsing via Parallel Token Prediction
by: Li, Lei, et al.
Published: (2026)
by: Li, Lei, et al.
Published: (2026)
Learnable Earth Parser: Discovering 3D Prototypes in Aerial Scans
by: Loiseau, Romain, et al.
Published: (2023)
by: Loiseau, Romain, et al.
Published: (2023)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
by: Lu, Jinghui, et al.
Published: (2024)
by: Lu, Jinghui, et al.
Published: (2024)
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
by: Nath, Oikantik, et al.
Published: (2025)
by: Nath, Oikantik, et al.
Published: (2025)
Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
by: Zhu, Fanwei, et al.
Published: (2025)
by: Zhu, Fanwei, et al.
Published: (2025)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
by: Luo, Chuwei, et al.
Published: (2024)
by: Luo, Chuwei, et al.
Published: (2024)
AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser
by: Ma, Ren, et al.
Published: (2025)
by: Ma, Ren, et al.
Published: (2025)
Tabular PDF Information Extraction with Local LLMs and Layout-Aware Parsing: A Reliability Evaluation
by: Hilmi, Muhammad Anis Al, et al.
Published: (2026)
by: Hilmi, Muhammad Anis Al, et al.
Published: (2026)
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
by: Li, Gengluo, et al.
Published: (2026)
by: Li, Gengluo, et al.
Published: (2026)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
by: Wang, Wenjie, et al.
Published: (2026)
by: Wang, Wenjie, et al.
Published: (2026)
ParseCaps: An Interpretable Parsing Capsule Network for Medical Image Diagnosis
by: Geng, Xinyu, et al.
Published: (2024)
by: Geng, Xinyu, et al.
Published: (2024)
DocParseNet: Advanced Semantic Segmentation and OCR Embeddings for Efficient Scanned Document Annotation
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
DocFusion: A Unified Framework for Document Parsing Tasks
by: Chai, Mingxu, et al.
Published: (2024)
by: Chai, Mingxu, et al.
Published: (2024)
Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model
by: Huang, Yaxuan, et al.
Published: (2025)
by: Huang, Yaxuan, et al.
Published: (2025)
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
by: An, Kaikai, et al.
Published: (2024)
by: An, Kaikai, et al.
Published: (2024)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
by: Li, Zhang, et al.
Published: (2025)
by: Li, Zhang, et al.
Published: (2025)
PointInfinity: Resolution-Invariant Point Diffusion Models
by: Huang, Zixuan, et al.
Published: (2024)
by: Huang, Zixuan, et al.
Published: (2024)
GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation
by: Li, Sifan, et al.
Published: (2026)
by: Li, Sifan, et al.
Published: (2026)
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
by: Feng, Hao, et al.
Published: (2025)
by: Feng, Hao, et al.
Published: (2025)
FineParser: A Fine-grained Spatio-temporal Action Parser for Human-centric Action Quality Assessment
by: Xu, Jinglin, et al.
Published: (2024)
by: Xu, Jinglin, et al.
Published: (2024)
Cross-domain Chinese Sentence Pattern Parsing
by: Yu, Jingsi, et al.
Published: (2024)
by: Yu, Jingsi, et al.
Published: (2024)
XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form Parser
by: Cheng, Xianfu, et al.
Published: (2024)
by: Cheng, Xianfu, et al.
Published: (2024)
Similar Items
-
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
by: Wang, Baode, et al.
Published: (2025) -
Parser-Oriented Structural Refinement for a Stable Layout Interface in Document Parsing
by: Liu, Fuyuan, et al.
Published: (2026) -
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
by: Jing, Hongyi, et al.
Published: (2025) -
UniParser: Multi-Human Parsing with Unified Correlation Representation Learning
by: Chu, Jiaming, et al.
Published: (2023) -
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding
by: Xu, Pengxin, et al.
Published: (2026)