Parser-Oriented Structural Refinement for a Stable Layout Interface in Document Parsing
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Fuyuan, Yu, Dianyu, Ren, He, Liu, Nayu, Kang, Xiaomian, Qiu, Delai, Zhang, Fa, Zhen, Genpeng, Liu, Shengping, Liang, Jiaen, Huang, Wei, Wang, Yining, Zhu, Junnan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PARL: Position-Aware Relation Learning Network for Document Layout Analysis
by: Liu, Fuyuan, et al.
Published: (2026)
by: Liu, Fuyuan, et al.
Published: (2026)
FocalOrder: Focal Preference Optimization for Reading Order Detection
by: Liu, Fuyuan, et al.
Published: (2026)
by: Liu, Fuyuan, et al.
Published: (2026)
Zipper-LoRA: Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition
by: Mei, Yuxiang, et al.
Published: (2026)
by: Mei, Yuxiang, et al.
Published: (2026)
LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verification
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
by: Wang, Baode, et al.
Published: (2025)
by: Wang, Baode, et al.
Published: (2025)
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
by: Wang, Baode, et al.
Published: (2025)
by: Wang, Baode, et al.
Published: (2025)
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning
by: Wei, Shuyu, et al.
Published: (2026)
by: Wei, Shuyu, et al.
Published: (2026)
Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
ASP2LJ : An Adversarial Self-Play Laywer Augmented Legal Judgment Framework
by: Chang, Ao, et al.
Published: (2025)
by: Chang, Ao, et al.
Published: (2025)
Semantic Pivots Enable Cross-Lingual Transfer in Large Language Models
by: He, Kaiyu, et al.
Published: (2025)
by: He, Kaiyu, et al.
Published: (2025)
Transparentize the Internal and External Knowledge Utilization in LLMs with Trustworthy Citation
by: Shen, Jiajun, et al.
Published: (2025)
by: Shen, Jiajun, et al.
Published: (2025)
LogParser-LLM: Advancing Efficient Log Parsing with Large Language Models
by: Zhong, Aoxiao, et al.
Published: (2024)
by: Zhong, Aoxiao, et al.
Published: (2024)
Optimizing Multi-Hop Document Retrieval Through Intermediate Representations
by: Lin, Jiaen, et al.
Published: (2025)
by: Lin, Jiaen, et al.
Published: (2025)
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers
by: Chen, Weidong, et al.
Published: (2026)
by: Chen, Weidong, et al.
Published: (2026)
dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
by: Li, Yumeng, et al.
Published: (2025)
by: Li, Yumeng, et al.
Published: (2025)
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding
by: Xu, Pengxin, et al.
Published: (2026)
by: Xu, Pengxin, et al.
Published: (2026)
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
by: Jing, Hongyi, et al.
Published: (2025)
by: Jing, Hongyi, et al.
Published: (2025)
UniParser: Multi-Human Parsing with Unified Correlation Representation Learning
by: Chu, Jiaming, et al.
Published: (2023)
by: Chu, Jiaming, et al.
Published: (2023)
N-gram Parsing for Jointly Training a Discriminative Constituency Parser
by: Arda Çelebi
Published: (2013)
by: Arda Çelebi
Published: (2013)
Refinement Module based on Parse Graph for Human Pose Estimation
by: Liu, Shibang, et al.
Published: (2025)
by: Liu, Shibang, et al.
Published: (2025)
VarParser: Unleashing the Neglected Power of Variables for LLM-based Log Parsing
by: Sun, Jinrui, et al.
Published: (2026)
by: Sun, Jinrui, et al.
Published: (2026)
Self-Modifying State Modeling for Simultaneous Machine Translation
by: Yu, Donglei, et al.
Published: (2024)
by: Yu, Donglei, et al.
Published: (2024)
How Much Can RAG Help the Reasoning of LLM?
by: Liu, Jingyu, et al.
Published: (2024)
by: Liu, Jingyu, et al.
Published: (2024)
Tackling the Inherent Difficulty of Noise Filtering in RAG
by: Liu, Jingyu, et al.
Published: (2026)
by: Liu, Jingyu, et al.
Published: (2026)
BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser
by: Ma, Ren, et al.
Published: (2025)
by: Ma, Ren, et al.
Published: (2025)
Cracking Factual Knowledge: A Comprehensive Analysis of Degenerate Knowledge Neurons in Large Language Models
by: Chen, Yuheng, et al.
Published: (2024)
by: Chen, Yuheng, et al.
Published: (2024)
How Do Document Parsers Break? Auditing Structural Vulnerability in Document Intelligence
by: Chen, Yue, et al.
Published: (2026)
by: Chen, Yue, et al.
Published: (2026)
Text Semantics to Flexible Design: A Residential Layout Generation Method Based on Stable Diffusion Model
by: Qiu, Zijin, et al.
Published: (2025)
by: Qiu, Zijin, et al.
Published: (2025)
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
by: Liu, Xiao, et al.
Published: (2026)
by: Liu, Xiao, et al.
Published: (2026)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
by: Li, Zhang, et al.
Published: (2025)
by: Li, Zhang, et al.
Published: (2025)
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
by: Nath, Oikantik, et al.
Published: (2025)
by: Nath, Oikantik, et al.
Published: (2025)
Solution to the 10th ABAW Expression Recognition Challenge: A Robust Multimodal Framework with Safe Cross-Attention and Modality Dropout
by: Yu, Jun, et al.
Published: (2026)
by: Yu, Jun, et al.
Published: (2026)
Sustainability management evolution: literature review and consolidative model
by: Ivete Delai
Published: (2016)
by: Ivete Delai
Published: (2016)
XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form Parser
by: Cheng, Xianfu, et al.
Published: (2024)
by: Cheng, Xianfu, et al.
Published: (2024)
LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
by: Heo, Inbum, et al.
Published: (2025)
by: Heo, Inbum, et al.
Published: (2025)
CREPE: Coordinate-Aware End-to-End Document Parser
by: Okamoto, Yamato, et al.
Published: (2024)
by: Okamoto, Yamato, et al.
Published: (2024)
Similar Items
-
PARL: Position-Aware Relation Learning Network for Document Layout Analysis
by: Liu, Fuyuan, et al.
Published: (2026) -
FocalOrder: Focal Preference Optimization for Reading Order Detection
by: Liu, Fuyuan, et al.
Published: (2026) -
Zipper-LoRA: Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition
by: Mei, Yuxiang, et al.
Published: (2026) -
LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verification
by: Li, Xuan, et al.
Published: (2026) -
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
by: Wang, Baode, et al.
Published: (2025)