GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Bakkali, Souhail, Biswas, Sanket, Ming, Zuheng, Coustaty, Mickaël, Rusiñol, Marçal, Terrades, Oriol Ramos, Lladós, Josep |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval
by: Molina, Adrià, et al.
Published: (2024)
by: Molina, Adrià, et al.
Published: (2024)
Visual Model Checking: Graph-Based Inference of Visual Routines for Image Retrieval
by: Molina, Adrià, et al.
Published: (2026)
by: Molina, Adrià, et al.
Published: (2026)
The OCR Quest for Generalization: Learning to recognize low-resource alphabets with model editing
by: Rodríguez, Adrià Molina, et al.
Published: (2025)
by: Rodríguez, Adrià Molina, et al.
Published: (2025)
IDTrust: Deep Identity Document Quality Detection with Bandpass Filtering
by: Al-Ghadi, Musab, et al.
Published: (2024)
by: Al-Ghadi, Musab, et al.
Published: (2024)
Hybrid Retrieval-Augmented Generation for Robust Multilingual Document Question Answering
by: Mudet, Anthony, et al.
Published: (2025)
by: Mudet, Anthony, et al.
Published: (2025)
Evaluating the Impact of Khmer Font Types on Text Recognition
by: Nom, Vannkinh, et al.
Published: (2025)
by: Nom, Vannkinh, et al.
Published: (2025)
KhmerST: A Low-Resource Khmer Scene Text Detection and Recognition Benchmark
by: Nom, Vannkinh, et al.
Published: (2024)
by: Nom, Vannkinh, et al.
Published: (2024)
LayeredDoc: Domain Adaptive Document Restoration with a Layer Separation Approach
by: Pilligua, Maria, et al.
Published: (2024)
by: Pilligua, Maria, et al.
Published: (2024)
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
by: Chau, Phan Phuong Mai, et al.
Published: (2024)
by: Chau, Phan Phuong Mai, et al.
Published: (2024)
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
by: Van Landeghem, Jordy, et al.
Published: (2024)
by: Van Landeghem, Jordy, et al.
Published: (2024)
SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition
by: Tiwari, Adarsh, et al.
Published: (2024)
by: Tiwari, Adarsh, et al.
Published: (2024)
GeoContrastNet: Contrastive Key-Value Edge Learning for Language-Agnostic Document Understanding
by: Biescas, Nil, et al.
Published: (2024)
by: Biescas, Nil, et al.
Published: (2024)
GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
DocSynthv2: A Practical Autoregressive Modeling for Document Generation
by: Biswas, Sanket, et al.
Published: (2024)
by: Biswas, Sanket, et al.
Published: (2024)
LLMChain: Blockchain-based Reputation System for Sharing and Evaluating Large Language Models
by: Bouchiha, Mouhamed Amine, et al.
Published: (2024)
by: Bouchiha, Mouhamed Amine, et al.
Published: (2024)
Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
by: Chattopadhyay, Soumitri, et al.
Published: (2024)
by: Chattopadhyay, Soumitri, et al.
Published: (2024)
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
Synthetic dataset of ID and Travel Document
by: Boned, Carlos, et al.
Published: (2024)
by: Boned, Carlos, et al.
Published: (2024)
DocRevive: A Unified Pipeline for Document Text Restoration
by: Purkayastha, Kunal, et al.
Published: (2026)
by: Purkayastha, Kunal, et al.
Published: (2026)
Recurrent Few-Shot model for Document Verification
by: Talarmain, Maxime, et al.
Published: (2024)
by: Talarmain, Maxime, et al.
Published: (2024)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
by: Das, Alloy, et al.
Published: (2024)
by: Das, Alloy, et al.
Published: (2024)
Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
by: Hamed, Omar, et al.
Published: (2024)
by: Hamed, Omar, et al.
Published: (2024)
The Role of Generative Systems in Historical Photography Management: A Case Study on Catalan Archives
by: Śanchez, Èric, et al.
Published: (2024)
by: Śanchez, Èric, et al.
Published: (2024)
Neurosymbolic Information Extraction from Transactional Documents
by: Hemmer, Arthur, et al.
Published: (2025)
by: Hemmer, Arthur, et al.
Published: (2025)
Confidence-Aware Document OCR Error Detection
by: Hemmer, Arthur, et al.
Published: (2024)
by: Hemmer, Arthur, et al.
Published: (2024)
La cadena global de valor en la industria electrónica
by: Josep Lladós Masllorens
Published: (2018)
by: Josep Lladós Masllorens
Published: (2018)
ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
by: Riera, Carlos Boned, et al.
Published: (2025)
by: Riera, Carlos Boned, et al.
Published: (2025)
LLM-Driven Medical Document Analysis: Enhancing Trustworthy Pathology and Differential Diagnosis
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
by: Jain, Chelsi, et al.
Published: (2025)
by: Jain, Chelsi, et al.
Published: (2025)
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
Structured Analysis and Comparison of Alphabets in Historical Handwritten Ciphers
by: Méndez, Martín, et al.
Published: (2024)
by: Méndez, Martín, et al.
Published: (2024)
CraftGraffiti: Exploring Human Identity with Custom Graffiti Art via Facial-Preserving Diffusion Models
by: Banerjee, Ayan, et al.
Published: (2025)
by: Banerjee, Ayan, et al.
Published: (2025)
LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
by: Kang, Zeyi, et al.
Published: (2025)
by: Kang, Zeyi, et al.
Published: (2025)
QUEST: Quality-aware Semi-supervised Table Extraction for Business Documents
by: Thomas, Eliott, et al.
Published: (2025)
by: Thomas, Eliott, et al.
Published: (2025)
Cross-Modal Adapter for Vision-Language Retrieval
by: Jiang, Haojun, et al.
Published: (2022)
by: Jiang, Haojun, et al.
Published: (2022)
DocReLM: Mastering Document Retrieval with Language Model
by: Wei, Gengchen, et al.
Published: (2024)
by: Wei, Gengchen, et al.
Published: (2024)
DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark
by: Hu, Ruofan, et al.
Published: (2026)
by: Hu, Ruofan, et al.
Published: (2026)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering
by: Banerjee, Ayan, et al.
Published: (2025)
by: Banerjee, Ayan, et al.
Published: (2025)
Similar Items
-
Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval
by: Molina, Adrià, et al.
Published: (2024) -
Visual Model Checking: Graph-Based Inference of Visual Routines for Image Retrieval
by: Molina, Adrià, et al.
Published: (2026) -
The OCR Quest for Generalization: Learning to recognize low-resource alphabets with model editing
by: Rodríguez, Adrià Molina, et al.
Published: (2025) -
IDTrust: Deep Identity Document Quality Detection with Bandpass Filtering
by: Al-Ghadi, Musab, et al.
Published: (2024) -
Hybrid Retrieval-Augmented Generation for Robust Multilingual Document Question Answering
by: Mudet, Anthony, et al.
Published: (2025)