ROAP: A Reading-Order and Attention-Prior Pipeline for Optimizing Layout Transformers in Key Information Extraction
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Tingwei, He, Jinxin, Song, Yonghong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
by: Tian, Yuanhe, et al.
Published: (2025)
by: Tian, Yuanhe, et al.
Published: (2025)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
by: Ji, Yifan, et al.
Published: (2026)
by: Ji, Yifan, et al.
Published: (2026)
Order Is Not Layout: Order-to-Space Bias in Image Generation
by: Zhang, Yongkang, et al.
Published: (2026)
by: Zhang, Yongkang, et al.
Published: (2026)
Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
by: Zhu, Fanwei, et al.
Published: (2025)
by: Zhu, Fanwei, et al.
Published: (2025)
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
by: Song, Yingjin, et al.
Published: (2025)
by: Song, Yingjin, et al.
Published: (2025)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
by: Luo, Chuwei, et al.
Published: (2024)
by: Luo, Chuwei, et al.
Published: (2024)
PharmaShip: An Entity-Centric, Reading-Order-Supervised Benchmark for Chinese Pharmaceutical Shipping Documents
by: Xie, Tingwei, et al.
Published: (2025)
by: Xie, Tingwei, et al.
Published: (2025)
RealKIE: Five Novel Datasets for Enterprise Key Information Extraction
by: Townsend, Benjamin, et al.
Published: (2024)
by: Townsend, Benjamin, et al.
Published: (2024)
Éclair -- Extracting Content and Layout with Integrated Reading Order for Documents
by: Karmanov, Ilia, et al.
Published: (2025)
by: Karmanov, Ilia, et al.
Published: (2025)
Co-Layout: LLM-driven Co-optimization for Interior Layout
by: Xiang, Chucheng, et al.
Published: (2025)
by: Xiang, Chucheng, et al.
Published: (2025)
Reading Order Independent Metrics for Information Extraction in Handwritten Documents
by: Villanova-Aparisi, David, et al.
Published: (2024)
by: Villanova-Aparisi, David, et al.
Published: (2024)
Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction
by: Pennec, Galann, et al.
Published: (2025)
by: Pennec, Galann, et al.
Published: (2025)
From Codicology to Code: A Comparative Study of Transformer and YOLO-based Detectors for Layout Analysis in Historical Documents
by: Aguilar, Sergio Torres
Published: (2025)
by: Aguilar, Sergio Torres
Published: (2025)
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
by: Heo, Inbum, et al.
Published: (2026)
by: Heo, Inbum, et al.
Published: (2026)
Joint Extraction Matters: Prompt-Based Visual Question Answering for Multi-Field Document Information Extraction
by: Loem, Mengsay, et al.
Published: (2025)
by: Loem, Mengsay, et al.
Published: (2025)
A U-Net and Transformer Pipeline for Multilingual Image Translation
by: Sahay, Siddharth, et al.
Published: (2025)
by: Sahay, Siddharth, et al.
Published: (2025)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
by: Mao, Weian, et al.
Published: (2026)
by: Mao, Weian, et al.
Published: (2026)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
by: Mao, Zhiming, et al.
Published: (2024)
by: Mao, Zhiming, et al.
Published: (2024)
Enhancing Document Key Information Localization Through Data Augmentation
by: Dai, Yue
Published: (2025)
by: Dai, Yue
Published: (2025)
Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
by: Duan, Changxu
Published: (2025)
by: Duan, Changxu
Published: (2025)
Optimizing Multimodal Language Models through Attention-based Interpretability
by: Sergeev, Alexander, et al.
Published: (2025)
by: Sergeev, Alexander, et al.
Published: (2025)
Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
by: Naz, Zubia, et al.
Published: (2025)
by: Naz, Zubia, et al.
Published: (2025)
FocalOrder: Focal Preference Optimization for Reading Order Detection
by: Liu, Fuyuan, et al.
Published: (2026)
by: Liu, Fuyuan, et al.
Published: (2026)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
by: Guo, Jialong, et al.
Published: (2024)
by: Guo, Jialong, et al.
Published: (2024)
XAttention: Block Sparse Attention with Antidiagonal Scoring
by: Xu, Ruyi, et al.
Published: (2025)
by: Xu, Ruyi, et al.
Published: (2025)
LAPDoc: Layout-Aware Prompting for Documents
by: Lamott, Marcel, et al.
Published: (2024)
by: Lamott, Marcel, et al.
Published: (2024)
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
SignDATA: Data Pipeline for Sign Language Translation
by: Chen, Kuanwei, et al.
Published: (2026)
by: Chen, Kuanwei, et al.
Published: (2026)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
by: Liang, Qiao, et al.
Published: (2025)
by: Liang, Qiao, et al.
Published: (2025)
Pose Priors from Language Models
by: Subramanian, Sanjay, et al.
Published: (2024)
by: Subramanian, Sanjay, et al.
Published: (2024)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
by: Gan, Chengguang, et al.
Published: (2025)
by: Gan, Chengguang, et al.
Published: (2025)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
Insight: A Multi-Modal Diagnostic Pipeline using LLMs for Ocular Surface Disease Diagnosis
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
by: Guan, Shuhao, et al.
Published: (2025)
by: Guan, Shuhao, et al.
Published: (2025)
Transformer with Controlled Attention for Synchronous Motion Captioning
by: Radouane, Karim, et al.
Published: (2024)
by: Radouane, Karim, et al.
Published: (2024)
Model Interpretability and Rationale Extraction by Input Mask Optimization
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
READoc: A Unified Benchmark for Realistic Document Structured Extraction
by: Li, Zichao, et al.
Published: (2024)
by: Li, Zichao, et al.
Published: (2024)
ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents
by: Pala, Furkan, et al.
Published: (2024)
by: Pala, Furkan, et al.
Published: (2024)
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Similar Items
-
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
by: Fan, Yue, et al.
Published: (2024) -
Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
by: Tian, Yuanhe, et al.
Published: (2025) -
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
by: Ji, Yifan, et al.
Published: (2026) -
Order Is Not Layout: Order-to-Space Bias in Image Generation
by: Zhang, Yongkang, et al.
Published: (2026) -
Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
by: Zhu, Fanwei, et al.
Published: (2025)