KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Heakl, Ahmed, Sohail, Abdullah, Ranjan, Mukul, Hossam, Rania, Ahmad, Ghazi Shazan, El-Geish, Mohamed, Maher, Omar, Shen, Zhiqiang, Khan, Fahad, Khan, Salman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
por: Ahmad, Ghazi Shazan, et al.
Publicado: (2025)
por: Ahmad, Ghazi Shazan, et al.
Publicado: (2025)
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
por: Ghaboura, Sara, et al.
Publicado: (2024)
por: Ghaboura, Sara, et al.
Publicado: (2024)
AIN: The Arabic INclusive Large Multimodal Model
por: Heakl, Ahmed, et al.
Publicado: (2025)
por: Heakl, Ahmed, et al.
Publicado: (2025)
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
por: Heakl, Ahmed, et al.
Publicado: (2026)
por: Heakl, Ahmed, et al.
Publicado: (2026)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
por: Heakl, Ahmed, et al.
Publicado: (2026)
por: Heakl, Ahmed, et al.
Publicado: (2026)
ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs
por: Heakl, Ahmed, et al.
Publicado: (2024)
por: Heakl, Ahmed, et al.
Publicado: (2024)
WorldCache: Content-Aware Caching for Accelerated Video World Models
por: Nawaz, Umair, et al.
Publicado: (2026)
por: Nawaz, Umair, et al.
Publicado: (2026)
Precision Aquaculture: An Integrated Computer Vision and IoT Approach for Optimized Tilapia Feeding
por: Hossam, Rania, et al.
Publicado: (2024)
por: Hossam, Rania, et al.
Publicado: (2024)
One Last Attention for Your Vision-Language Model
por: Chen, Liang, et al.
Publicado: (2025)
por: Chen, Liang, et al.
Publicado: (2025)
Planner and Executor: Collaboration between Discrete Diffusion And Autoregressive Models in Reasoning
por: Berrayana, Lina, et al.
Publicado: (2025)
por: Berrayana, Lina, et al.
Publicado: (2025)
Language Guided Domain Generalized Medical Image Segmentation
por: Kunhimon, Shahina, et al.
Publicado: (2024)
por: Kunhimon, Shahina, et al.
Publicado: (2024)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
por: Bharadwaj, Rohit, et al.
Publicado: (2024)
por: Bharadwaj, Rohit, et al.
Publicado: (2024)
Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
por: Ali, Muhammad, et al.
Publicado: (2025)
por: Ali, Muhammad, et al.
Publicado: (2025)
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding
por: Patle, Shubham, et al.
Publicado: (2026)
por: Patle, Shubham, et al.
Publicado: (2026)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
por: Mahmood, Ahmad, et al.
Publicado: (2024)
por: Mahmood, Ahmad, et al.
Publicado: (2024)
ArabicNumBench: Evaluating Arabic Number Reading in Large Language Models
por: Alhumud, Anas, et al.
Publicado: (2026)
por: Alhumud, Anas, et al.
Publicado: (2026)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
por: Maaz, Muhammad, et al.
Publicado: (2024)
por: Maaz, Muhammad, et al.
Publicado: (2024)
CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark
por: Heakl, Ahmed, et al.
Publicado: (2025)
por: Heakl, Ahmed, et al.
Publicado: (2025)
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
por: Ghaboura, Sara, et al.
Publicado: (2025)
por: Ghaboura, Sara, et al.
Publicado: (2025)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
por: Chen, Shiming, et al.
Publicado: (2025)
por: Chen, Shiming, et al.
Publicado: (2025)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
por: Dharmasiri, Amaya, et al.
Publicado: (2024)
por: Dharmasiri, Amaya, et al.
Publicado: (2024)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
por: Maaz, Muhammad, et al.
Publicado: (2023)
por: Maaz, Muhammad, et al.
Publicado: (2023)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
por: Maaz, Muhammad, et al.
Publicado: (2025)
por: Maaz, Muhammad, et al.
Publicado: (2025)
MedContext: Learning Contextual Cues for Efficient Volumetric Medical Segmentation
por: Gani, Hanan, et al.
Publicado: (2024)
por: Gani, Hanan, et al.
Publicado: (2024)
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
por: Hennara, Khalil, et al.
Publicado: (2025)
por: Hennara, Khalil, et al.
Publicado: (2025)
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework
por: Kumar, Komal, et al.
Publicado: (2026)
por: Kumar, Komal, et al.
Publicado: (2026)
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
por: Malik, Hashmat Shadab, et al.
Publicado: (2025)
por: Malik, Hashmat Shadab, et al.
Publicado: (2025)
On the Cultural Anachronism and Temporal Reasoning in Vision Language Models
por: Ranjan, Mukul, et al.
Publicado: (2026)
por: Ranjan, Mukul, et al.
Publicado: (2026)
AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
por: Nawaz, Umair, et al.
Publicado: (2024)
por: Nawaz, Umair, et al.
Publicado: (2024)
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
por: Shabbir, Akashah, et al.
Publicado: (2025)
por: Shabbir, Akashah, et al.
Publicado: (2025)
Dr.LLM: Dynamic Layer Routing in LLMs
por: Heakl, Ahmed, et al.
Publicado: (2025)
por: Heakl, Ahmed, et al.
Publicado: (2025)
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
por: Chen, Shiming, et al.
Publicado: (2025)
por: Chen, Shiming, et al.
Publicado: (2025)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
por: Chen, Shiming, et al.
Publicado: (2024)
por: Chen, Shiming, et al.
Publicado: (2024)
Towards Evaluating the Robustness of Visual State Space Models
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
PEACH: A sentence-aligned Parallel English-Arabic Corpus for Healthcare
por: Al-Sabbagh, Rania
Publicado: (2025)
por: Al-Sabbagh, Rania
Publicado: (2025)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
por: Demidov, Dmitry, et al.
Publicado: (2025)
por: Demidov, Dmitry, et al.
Publicado: (2025)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
por: Noman, Mubashir, et al.
Publicado: (2024)
por: Noman, Mubashir, et al.
Publicado: (2024)
Arabic Handwritten Text for Person Biometric Identification: A Deep Learning Approach
por: Balat, Mazen, et al.
Publicado: (2024)
por: Balat, Mazen, et al.
Publicado: (2024)
Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction
por: Rashad, Mohamed
Publicado: (2024)
por: Rashad, Mohamed
Publicado: (2024)
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
por: Al-Homoud, Haneen, et al.
Publicado: (2025)
por: Al-Homoud, Haneen, et al.
Publicado: (2025)
Ejemplares similares
-
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
por: Ahmad, Ghazi Shazan, et al.
Publicado: (2025) -
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
por: Ghaboura, Sara, et al.
Publicado: (2024) -
AIN: The Arabic INclusive Large Multimodal Model
por: Heakl, Ahmed, et al.
Publicado: (2025) -
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
por: Heakl, Ahmed, et al.
Publicado: (2026) -
DocAtlas: Multilingual Document Understanding Across 80+ Languages
por: Heakl, Ahmed, et al.
Publicado: (2026)