DocAtlas: Multilingual Document Understanding Across 80+ Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Heakl, Ahmed, Mohamed, Youssef, Sohail, Abdullah, Elbadry, Rania, Nassar, Ahmed, Staar, Peter W. J., Khan, Fahad Shahbaz, Razzak, Imran, Khan, Salman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
by: Heakl, Ahmed, et al.
Published: (2026)
by: Heakl, Ahmed, et al.
Published: (2026)
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
WorldCache: Content-Aware Caching for Accelerated Video World Models
by: Nawaz, Umair, et al.
Published: (2026)
by: Nawaz, Umair, et al.
Published: (2026)
AIN: The Arabic INclusive Large Multimodal Model
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models
by: Heakl, Ahmed, et al.
Published: (2024)
by: Heakl, Ahmed, et al.
Published: (2024)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
Planner and Executor: Collaboration between Discrete Diffusion And Autoregressive Models in Reasoning
by: Berrayana, Lina, et al.
Published: (2025)
by: Berrayana, Lina, et al.
Published: (2025)
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
by: Deria, Ankan, et al.
Published: (2026)
by: Deria, Ankan, et al.
Published: (2026)
Advancing Ear Biometrics: Enhancing Accuracy and Robustness through Deep Learning
by: Mohamed, Youssef, et al.
Published: (2024)
by: Mohamed, Youssef, et al.
Published: (2024)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
by: Maaz, Muhammad, et al.
Published: (2023)
by: Maaz, Muhammad, et al.
Published: (2023)
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
by: Shaker, Abdelrahman, et al.
Published: (2026)
by: Shaker, Abdelrahman, et al.
Published: (2026)
Arabic Handwritten Text for Person Biometric Identification: A Deep Learning Approach
by: Balat, Mazen, et al.
Published: (2024)
by: Balat, Mazen, et al.
Published: (2024)
ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs
by: Heakl, Ahmed, et al.
Published: (2024)
by: Heakl, Ahmed, et al.
Published: (2024)
AraSpider: Democratizing Arabic-to-SQL
by: Heakl, Ahmed, et al.
Published: (2024)
by: Heakl, Ahmed, et al.
Published: (2024)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
by: Mahmood, Ahmad, et al.
Published: (2024)
by: Mahmood, Ahmad, et al.
Published: (2024)
Language Guided Domain Generalized Medical Image Segmentation
by: Kunhimon, Shahina, et al.
Published: (2024)
by: Kunhimon, Shahina, et al.
Published: (2024)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
by: Chen, Shiming, et al.
Published: (2025)
by: Chen, Shiming, et al.
Published: (2025)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
by: Dharmasiri, Amaya, et al.
Published: (2024)
by: Dharmasiri, Amaya, et al.
Published: (2024)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
by: Maaz, Muhammad, et al.
Published: (2025)
by: Maaz, Muhammad, et al.
Published: (2025)
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
by: Chen, Shiming, et al.
Published: (2025)
by: Chen, Shiming, et al.
Published: (2025)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
by: Chen, Shiming, et al.
Published: (2024)
by: Chen, Shiming, et al.
Published: (2024)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
by: Shaker, Abdelrahman, et al.
Published: (2025)
by: Shaker, Abdelrahman, et al.
Published: (2025)
Precision Aquaculture: An Integrated Computer Vision and IoT Approach for Optimized Tilapia Feeding
by: Hossam, Rania, et al.
Published: (2024)
by: Hossam, Rania, et al.
Published: (2024)
Enhancing Novel Object Detection via Cooperative Foundational Models
by: Bharadwaj, Rohit, et al.
Published: (2023)
by: Bharadwaj, Rohit, et al.
Published: (2023)
The Role of AI in Early Detection of Life-Threatening Diseases: A Retinal Imaging Perspective
by: Khan, Tariq M, et al.
Published: (2025)
by: Khan, Tariq M, et al.
Published: (2025)
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework
by: Kumar, Komal, et al.
Published: (2026)
by: Kumar, Komal, et al.
Published: (2026)
Tracking Meets Large Multimodal Models for Driving Scenario Understanding
by: Ishaq, Ayesha, et al.
Published: (2025)
by: Ishaq, Ayesha, et al.
Published: (2025)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
by: Demidov, Dmitry, et al.
Published: (2025)
by: Demidov, Dmitry, et al.
Published: (2025)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
by: Bharadwaj, Rohit, et al.
Published: (2024)
by: Bharadwaj, Rohit, et al.
Published: (2024)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
by: Ghaboura, Sara, et al.
Published: (2024)
by: Ghaboura, Sara, et al.
Published: (2024)
Dr.LLM: Dynamic Layer Routing in LLMs
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
by: Maaz, Muhammad, et al.
Published: (2024)
by: Maaz, Muhammad, et al.
Published: (2024)
MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
by: Sheikh, Tooba Tehreem, et al.
Published: (2025)
by: Sheikh, Tooba Tehreem, et al.
Published: (2025)
Towards Evaluating the Robustness of Visual State Space Models
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
by: Dissanayake, Dinura, et al.
Published: (2025)
by: Dissanayake, Dinura, et al.
Published: (2025)
Latent-DARM: Bridging Discrete Diffusion And Autoregressive Models For Reasoning
by: Berrayana, Lina, et al.
Published: (2026)
by: Berrayana, Lina, et al.
Published: (2026)
MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images
by: Deria, Ankan, et al.
Published: (2026)
by: Deria, Ankan, et al.
Published: (2026)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
by: Rasheed, Hanoona, et al.
Published: (2025)
by: Rasheed, Hanoona, et al.
Published: (2025)
Learnable Weight Initialization for Volumetric Medical Image Segmentation
by: Kunhimon, Shahina, et al.
Published: (2023)
by: Kunhimon, Shahina, et al.
Published: (2023)
Similar Items
-
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
by: Heakl, Ahmed, et al.
Published: (2026) -
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
by: Heakl, Ahmed, et al.
Published: (2025) -
WorldCache: Content-Aware Caching for Accelerated Video World Models
by: Nawaz, Umair, et al.
Published: (2026) -
AIN: The Arabic INclusive Large Multimodal Model
by: Heakl, Ahmed, et al.
Published: (2025) -
ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models
by: Heakl, Ahmed, et al.
Published: (2024)