DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | SR, Nikitha, Menta, Tarun Ram, Sarkar, Mausoom |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs
von: SR, Nikitha, et al.
Veröffentlicht: (2025)
von: SR, Nikitha, et al.
Veröffentlicht: (2025)
DesignSense: A Human Preference Dataset and Reward Modeling Framework for Graphic Layout Generation
von: Gopal, Varun, et al.
Veröffentlicht: (2026)
von: Gopal, Varun, et al.
Veröffentlicht: (2026)
Evaluating Variance in Visual Question Answering Benchmarks
von: SR, Nikitha
Veröffentlicht: (2025)
von: SR, Nikitha
Veröffentlicht: (2025)
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
von: Patnaik, Sohan, et al.
Veröffentlicht: (2025)
von: Patnaik, Sohan, et al.
Veröffentlicht: (2025)
EOPose : Exemplar-based object reposing using Generalized Pose Correspondences
von: Mehrotra, Sarthak, et al.
Veröffentlicht: (2025)
von: Mehrotra, Sarthak, et al.
Veröffentlicht: (2025)
AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
von: Anand, Neeraj, et al.
Veröffentlicht: (2025)
von: Anand, Neeraj, et al.
Veröffentlicht: (2025)
ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models
von: Srivastava, Ashutosh, et al.
Veröffentlicht: (2024)
von: Srivastava, Ashutosh, et al.
Veröffentlicht: (2024)
Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
von: Jadhav, Avadhoot, et al.
Veröffentlicht: (2025)
von: Jadhav, Avadhoot, et al.
Veröffentlicht: (2025)
LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
von: Heo, Inbum, et al.
Veröffentlicht: (2025)
von: Heo, Inbum, et al.
Veröffentlicht: (2025)
Diachronic Document Dataset for Semantic Layout Analysis
von: Clérice, Thibault, et al.
Veröffentlicht: (2024)
von: Clérice, Thibault, et al.
Veröffentlicht: (2024)
SFDLA: Source-Free Document Layout Analysis
von: Tewes, Sebastian, et al.
Veröffentlicht: (2025)
von: Tewes, Sebastian, et al.
Veröffentlicht: (2025)
A Hybrid Approach for Document Layout Analysis in Document images
von: Shehzadi, Tahira, et al.
Veröffentlicht: (2024)
von: Shehzadi, Tahira, et al.
Veröffentlicht: (2024)
Cross-Domain Document Layout Analysis Using Document Style Guide
von: Wu, Xingjiao, et al.
Veröffentlicht: (2022)
von: Wu, Xingjiao, et al.
Veröffentlicht: (2022)
Beyond Bag-of-Patches: Learning Global Layout via Textual Supervision for Late-Interaction Visual Document Retrieval
von: Tilli, Pascal, et al.
Veröffentlicht: (2026)
von: Tilli, Pascal, et al.
Veröffentlicht: (2026)
DLAFormer: An End-to-End Transformer For Document Layout Analysis
von: Wang, Jiawei, et al.
Veröffentlicht: (2024)
von: Wang, Jiawei, et al.
Veröffentlicht: (2024)
HybriDLA: Hybrid Generation for Document Layout Analysis
von: Chen, Yufan, et al.
Veröffentlicht: (2025)
von: Chen, Yufan, et al.
Veröffentlicht: (2025)
Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis
von: Bi, Tianci, et al.
Veröffentlicht: (2024)
von: Bi, Tianci, et al.
Veröffentlicht: (2024)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)
TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
von: Cao, Bingyi, et al.
Veröffentlicht: (2026)
von: Cao, Bingyi, et al.
Veröffentlicht: (2026)
UnSupDLA: Towards Unsupervised Document Layout Analysis
von: Sheikh, Talha Uddin, et al.
Veröffentlicht: (2024)
von: Sheikh, Talha Uddin, et al.
Veröffentlicht: (2024)
RoDLA: Benchmarking the Robustness of Document Layout Analysis Models
von: Chen, Yufan, et al.
Veröffentlicht: (2024)
von: Chen, Yufan, et al.
Veröffentlicht: (2024)
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
LLM-Guided Probabilistic Fusion for Label-Efficient Document Layout Analysis
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2025)
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2025)
PARL: Position-Aware Relation Learning Network for Document Layout Analysis
von: Liu, Fuyuan, et al.
Veröffentlicht: (2026)
von: Liu, Fuyuan, et al.
Veröffentlicht: (2026)
Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement
von: Guo, Junrong, et al.
Veröffentlicht: (2026)
von: Guo, Junrong, et al.
Veröffentlicht: (2026)
SciPostLayout: A Dataset for Layout Analysis and Layout Generation of Scientific Posters
von: Tanaka, Shohei, et al.
Veröffentlicht: (2024)
von: Tanaka, Shohei, et al.
Veröffentlicht: (2024)
Generating Animated Layouts as Structured Text Representations
von: Shin, Yeonsang, et al.
Veröffentlicht: (2025)
von: Shin, Yeonsang, et al.
Veröffentlicht: (2025)
The COTe score: A decomposable framework for evaluating Document Layout Analysis models
von: Bourne, Jonathan, et al.
Veröffentlicht: (2026)
von: Bourne, Jonathan, et al.
Veröffentlicht: (2026)
HandDreamer: Zero-Shot Text to 3D Hand Model Generation using Corrective Hand Shape Guidance
von: Rosh, Green, et al.
Veröffentlicht: (2026)
von: Rosh, Green, et al.
Veröffentlicht: (2026)
Towards Khmer Scene Document Layout Detection
von: Kong, Marry, et al.
Veröffentlicht: (2026)
von: Kong, Marry, et al.
Veröffentlicht: (2026)
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training
von: Jiang, Zhouqiang, et al.
Veröffentlicht: (2024)
von: Jiang, Zhouqiang, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
von: Gan, Lubin, et al.
Veröffentlicht: (2025)
von: Gan, Lubin, et al.
Veröffentlicht: (2025)
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
von: Heo, Inbum, et al.
Veröffentlicht: (2026)
von: Heo, Inbum, et al.
Veröffentlicht: (2026)
Accurate Fine-grained Layout Analysis for the Historical Tibetan Document Based on the Instance Segmentation
von: Zhao, Penghai, et al.
Veröffentlicht: (2021)
von: Zhao, Penghai, et al.
Veröffentlicht: (2021)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
Multitwine: Multi-Object Compositing with Text and Layout Control
von: Tarrés, Gemma Canet, et al.
Veröffentlicht: (2025)
von: Tarrés, Gemma Canet, et al.
Veröffentlicht: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
von: Li, Yuheng, et al.
Veröffentlicht: (2024)
von: Li, Yuheng, et al.
Veröffentlicht: (2024)
Scan-and-Print: Patch-level Data Summarization and Augmentation for Content-aware Layout Generation in Poster Design
von: Hsu, HsiaoYuan, et al.
Veröffentlicht: (2025)
von: Hsu, HsiaoYuan, et al.
Veröffentlicht: (2025)
Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long Text
von: Liang, Guotao, et al.
Veröffentlicht: (2025)
von: Liang, Guotao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs
von: SR, Nikitha, et al.
Veröffentlicht: (2025) -
DesignSense: A Human Preference Dataset and Reward Modeling Framework for Graphic Layout Generation
von: Gopal, Varun, et al.
Veröffentlicht: (2026) -
Evaluating Variance in Visual Question Answering Benchmarks
von: SR, Nikitha
Veröffentlicht: (2025) -
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
von: Patnaik, Sohan, et al.
Veröffentlicht: (2025) -
EOPose : Exemplar-based object reposing using Generalized Pose Correspondences
von: Mehrotra, Sarthak, et al.
Veröffentlicht: (2025)