ROAP: A Reading-Order and Attention-Prior Pipeline for Optimizing Layout Transformers in Key Information Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Tingwei, He, Jinxin, Song, Yonghong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
Order Is Not Layout: Order-to-Space Bias in Image Generation
von: Zhang, Yongkang, et al.
Veröffentlicht: (2026)
von: Zhang, Yongkang, et al.
Veröffentlicht: (2026)
Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
von: Zhu, Fanwei, et al.
Veröffentlicht: (2025)
von: Zhu, Fanwei, et al.
Veröffentlicht: (2025)
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
von: Song, Yingjin, et al.
Veröffentlicht: (2025)
von: Song, Yingjin, et al.
Veröffentlicht: (2025)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
PharmaShip: An Entity-Centric, Reading-Order-Supervised Benchmark for Chinese Pharmaceutical Shipping Documents
von: Xie, Tingwei, et al.
Veröffentlicht: (2025)
von: Xie, Tingwei, et al.
Veröffentlicht: (2025)
RealKIE: Five Novel Datasets for Enterprise Key Information Extraction
von: Townsend, Benjamin, et al.
Veröffentlicht: (2024)
von: Townsend, Benjamin, et al.
Veröffentlicht: (2024)
Éclair -- Extracting Content and Layout with Integrated Reading Order for Documents
von: Karmanov, Ilia, et al.
Veröffentlicht: (2025)
von: Karmanov, Ilia, et al.
Veröffentlicht: (2025)
Co-Layout: LLM-driven Co-optimization for Interior Layout
von: Xiang, Chucheng, et al.
Veröffentlicht: (2025)
von: Xiang, Chucheng, et al.
Veröffentlicht: (2025)
Reading Order Independent Metrics for Information Extraction in Handwritten Documents
von: Villanova-Aparisi, David, et al.
Veröffentlicht: (2024)
von: Villanova-Aparisi, David, et al.
Veröffentlicht: (2024)
Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction
von: Pennec, Galann, et al.
Veröffentlicht: (2025)
von: Pennec, Galann, et al.
Veröffentlicht: (2025)
From Codicology to Code: A Comparative Study of Transformer and YOLO-based Detectors for Layout Analysis in Historical Documents
von: Aguilar, Sergio Torres
Veröffentlicht: (2025)
von: Aguilar, Sergio Torres
Veröffentlicht: (2025)
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
von: Heo, Inbum, et al.
Veröffentlicht: (2026)
von: Heo, Inbum, et al.
Veröffentlicht: (2026)
Joint Extraction Matters: Prompt-Based Visual Question Answering for Multi-Field Document Information Extraction
von: Loem, Mengsay, et al.
Veröffentlicht: (2025)
von: Loem, Mengsay, et al.
Veröffentlicht: (2025)
A U-Net and Transformer Pipeline for Multilingual Image Translation
von: Sahay, Siddharth, et al.
Veröffentlicht: (2025)
von: Sahay, Siddharth, et al.
Veröffentlicht: (2025)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
von: Mao, Weian, et al.
Veröffentlicht: (2026)
von: Mao, Weian, et al.
Veröffentlicht: (2026)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)
Enhancing Document Key Information Localization Through Data Augmentation
von: Dai, Yue
Veröffentlicht: (2025)
von: Dai, Yue
Veröffentlicht: (2025)
Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
von: Duan, Changxu
Veröffentlicht: (2025)
von: Duan, Changxu
Veröffentlicht: (2025)
Optimizing Multimodal Language Models through Attention-based Interpretability
von: Sergeev, Alexander, et al.
Veröffentlicht: (2025)
von: Sergeev, Alexander, et al.
Veröffentlicht: (2025)
Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
von: Naz, Zubia, et al.
Veröffentlicht: (2025)
von: Naz, Zubia, et al.
Veröffentlicht: (2025)
FocalOrder: Focal Preference Optimization for Reading Order Detection
von: Liu, Fuyuan, et al.
Veröffentlicht: (2026)
von: Liu, Fuyuan, et al.
Veröffentlicht: (2026)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
von: Guo, Jialong, et al.
Veröffentlicht: (2024)
von: Guo, Jialong, et al.
Veröffentlicht: (2024)
XAttention: Block Sparse Attention with Antidiagonal Scoring
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
LAPDoc: Layout-Aware Prompting for Documents
von: Lamott, Marcel, et al.
Veröffentlicht: (2024)
von: Lamott, Marcel, et al.
Veröffentlicht: (2024)
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
SignDATA: Data Pipeline for Sign Language Translation
von: Chen, Kuanwei, et al.
Veröffentlicht: (2026)
von: Chen, Kuanwei, et al.
Veröffentlicht: (2026)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
Pose Priors from Language Models
von: Subramanian, Sanjay, et al.
Veröffentlicht: (2024)
von: Subramanian, Sanjay, et al.
Veröffentlicht: (2024)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Insight: A Multi-Modal Diagnostic Pipeline using LLMs for Ocular Surface Disease Diagnosis
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2024)
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2024)
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
von: Guan, Shuhao, et al.
Veröffentlicht: (2025)
von: Guan, Shuhao, et al.
Veröffentlicht: (2025)
Transformer with Controlled Attention for Synchronous Motion Captioning
von: Radouane, Karim, et al.
Veröffentlicht: (2024)
von: Radouane, Karim, et al.
Veröffentlicht: (2024)
Model Interpretability and Rationale Extraction by Input Mask Optimization
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
READoc: A Unified Benchmark for Realistic Document Structured Extraction
von: Li, Zichao, et al.
Veröffentlicht: (2024)
von: Li, Zichao, et al.
Veröffentlicht: (2024)
ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
von: Fan, Yue, et al.
Veröffentlicht: (2024) -
Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025) -
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
von: Ji, Yifan, et al.
Veröffentlicht: (2026) -
Order Is Not Layout: Order-to-Space Bias in Image Generation
von: Zhang, Yongkang, et al.
Veröffentlicht: (2026) -
Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
von: Zhu, Fanwei, et al.
Veröffentlicht: (2025)