PILOT: A Promptable Interleaved Layout-aware OCR Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Hamdi, Laziz, Tamasna, Amine, Boisson, Pascal, Paquet, Thierry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TableSeq: Unified Generation of Structure, Content, and Layout
by: Hamdi, Laziz, et al.
Published: (2026)
by: Hamdi, Laziz, et al.
Published: (2026)
FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers
by: Hamdi, Laziz, et al.
Published: (2026)
by: Hamdi, Laziz, et al.
Published: (2026)
DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates
by: Hamdi, Laziz, et al.
Published: (2026)
by: Hamdi, Laziz, et al.
Published: (2026)
A Benchmark of State-Space Models vs. Transformers and BiLSTM-based Models for Historical Newspaper OCR
by: Agbeti-messan, Merveilles, et al.
Published: (2026)
by: Agbeti-messan, Merveilles, et al.
Published: (2026)
DianJin-OCR-R1: Enhancing OCR Capabilities via a Reasoning-and-Tool Interleaved Vision-Language Model
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
Keypoint Promptable Re-Identification
by: Somers, Vladimir, et al.
Published: (2024)
by: Somers, Vladimir, et al.
Published: (2024)
Loom: Diffusion-Transformer for Interleaved Generation
by: Ye, Mingcheng, et al.
Published: (2025)
by: Ye, Mingcheng, et al.
Published: (2025)
Revisiting Transformers with Insights from Image Filtering and Boosting
by: Abdullaev, Laziz U., et al.
Published: (2025)
by: Abdullaev, Laziz U., et al.
Published: (2025)
PHAC: Promptable Human Amodal Completion
by: Noh, Seung Young, et al.
Published: (2026)
by: Noh, Seung Young, et al.
Published: (2026)
Robust Promptable Video Object Segmentation
by: Lee, Sohyun, et al.
Published: (2026)
by: Lee, Sohyun, et al.
Published: (2026)
Layout Anything: One Transformer for Universal Room Layout Estimation
by: Mia, Md Sohag, et al.
Published: (2025)
by: Mia, Md Sohag, et al.
Published: (2025)
Retrieval-Augmented Layout Transformer for Content-Aware Layout Generation
by: Horita, Daichi, et al.
Published: (2023)
by: Horita, Daichi, et al.
Published: (2023)
Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription
by: Ríos-Vila, Antonio, et al.
Published: (2024)
by: Ríos-Vila, Antonio, et al.
Published: (2024)
LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer
by: Yu, Ning, et al.
Published: (2022)
by: Yu, Ning, et al.
Published: (2022)
PromptHMR: Promptable Human Mesh Recovery
by: Wang, Yufu, et al.
Published: (2025)
by: Wang, Yufu, et al.
Published: (2025)
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
by: Zhang, Hui, et al.
Published: (2024)
by: Zhang, Hui, et al.
Published: (2024)
LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
Classifying the Unknown: In-Context Learning for Open-Vocabulary Text and Symbol Recognition
by: Simon, Tom, et al.
Published: (2025)
by: Simon, Tom, et al.
Published: (2025)
End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music
by: Ríos-Vila, Antonio, et al.
Published: (2024)
by: Ríos-Vila, Antonio, et al.
Published: (2024)
ERPA: Efficient RPA Model Integrating OCR and LLMs for Intelligent Document Processing
by: Abdellaif, Osama, et al.
Published: (2024)
by: Abdellaif, Osama, et al.
Published: (2024)
Transferring to Real-World Layouts: A Depth-aware Framework for Scene Adaptation
by: Chen, Mu, et al.
Published: (2023)
by: Chen, Mu, et al.
Published: (2023)
nnInteractive: Redefining 3D Promptable Segmentation
by: Isensee, Fabian, et al.
Published: (2025)
by: Isensee, Fabian, et al.
Published: (2025)
PRISM: A Promptable and Robust Interactive Segmentation Model with Visual Prompts
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Social-Transmotion: Promptable Human Trajectory Prediction
by: Saadatnejad, Saeed, et al.
Published: (2023)
by: Saadatnejad, Saeed, et al.
Published: (2023)
Text-Promptable Propagation for Referring Medical Image Sequence Segmentation
by: Yuan, Runtian, et al.
Published: (2025)
by: Yuan, Runtian, et al.
Published: (2025)
Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning
by: Yang, Hui-Yue, et al.
Published: (2024)
by: Yang, Hui-Yue, et al.
Published: (2024)
Leveraging Hallucinations to Reduce Manual Prompt Dependency in Promptable Segmentation
by: Hu, Jian, et al.
Published: (2024)
by: Hu, Jian, et al.
Published: (2024)
Rethinking Text-Promptable Surgical Instrument Segmentation with Robust Framework
by: Choi, Tae-Min, et al.
Published: (2024)
by: Choi, Tae-Min, et al.
Published: (2024)
PROMO: Promptable Outfitting for Efficient High-Fidelity Virtual Try-On
by: Chen, Haohua, et al.
Published: (2026)
by: Chen, Haohua, et al.
Published: (2026)
Text Promptable Surgical Instrument Segmentation with Vision-Language Models
by: Zhou, Zijian, et al.
Published: (2023)
by: Zhou, Zijian, et al.
Published: (2023)
Locate, Assign, Refine: Taming Customized Promptable Image Inpainting
by: Pan, Yulin, et al.
Published: (2024)
by: Pan, Yulin, et al.
Published: (2024)
Beyond Bag-of-Patches: Learning Global Layout via Textual Supervision for Late-Interaction Visual Document Retrieval
by: Tilli, Pascal, et al.
Published: (2026)
by: Tilli, Pascal, et al.
Published: (2026)
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940
by: Constum, Thomas, et al.
Published: (2024)
by: Constum, Thomas, et al.
Published: (2024)
DLAFormer: An End-to-End Transformer For Document Layout Analysis
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
OCR-Agent: Agentic OCR with Capability and Memory Reflection
by: Wen, Shimin, et al.
Published: (2026)
by: Wen, Shimin, et al.
Published: (2026)
OmniOCR: Generalist OCR for Ethnic Minority Languages
by: Liu, Bonan, et al.
Published: (2026)
by: Liu, Bonan, et al.
Published: (2026)
SegSLR: Promptable Video Segmentation for Isolated Sign Language Recognition
by: Schreiber, Sven, et al.
Published: (2025)
by: Schreiber, Sven, et al.
Published: (2025)
INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
MV-SAM: Multi-view Promptable Segmentation using Pointmap Guidance
by: Jeong, Yoonwoo, et al.
Published: (2026)
by: Jeong, Yoonwoo, et al.
Published: (2026)
Similar Items
-
TableSeq: Unified Generation of Structure, Content, and Layout
by: Hamdi, Laziz, et al.
Published: (2026) -
FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers
by: Hamdi, Laziz, et al.
Published: (2026) -
DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates
by: Hamdi, Laziz, et al.
Published: (2026) -
A Benchmark of State-Space Models vs. Transformers and BiLSTM-based Models for Historical Newspaper OCR
by: Agbeti-messan, Merveilles, et al.
Published: (2026) -
DianJin-OCR-R1: Enhancing OCR Capabilities via a Reasoning-and-Tool Interleaved Vision-Language Model
by: Chen, Qian, et al.
Published: (2025)