HalalBench: A Multilingual OCR Benchmark for Food Packaging Ingredient Extraction
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Arief, Hasan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation
von: Yang, Qi, et al.
Veröffentlicht: (2026)
von: Yang, Qi, et al.
Veröffentlicht: (2026)
Enhancing OCR for Sino-Vietnamese Language Processing via Fine-tuned PaddleOCRv5
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2025)
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2025)
Large Language Models for Simultaneous Named Entity Extraction and Spelling Correction
von: Whittaker, Edward, et al.
Veröffentlicht: (2024)
von: Whittaker, Edward, et al.
Veröffentlicht: (2024)
A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports
von: Schäfer, Henning, et al.
Veröffentlicht: (2025)
von: Schäfer, Henning, et al.
Veröffentlicht: (2025)
Devanagari Handwritten Character Recognition using Convolutional Neural Network
von: Mehta, Diksha, et al.
Veröffentlicht: (2025)
von: Mehta, Diksha, et al.
Veröffentlicht: (2025)
Document Understanding for Healthcare Referrals
von: Mistry, Jimit, et al.
Veröffentlicht: (2023)
von: Mistry, Jimit, et al.
Veröffentlicht: (2023)
An accurate and revised version of optical character recognition-based speech synthesis using LabVIEW
von: Mehta, Prateek, et al.
Veröffentlicht: (2025)
von: Mehta, Prateek, et al.
Veröffentlicht: (2025)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
von: Zmanovskii, Nikita
Veröffentlicht: (2025)
von: Zmanovskii, Nikita
Veröffentlicht: (2025)
Hydra: Unifying Document Retrieval and Generation in a Single Vision-Language Model
von: Georgiou, Athos
Veröffentlicht: (2026)
von: Georgiou, Athos
Veröffentlicht: (2026)
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
von: Deichler, Anna, et al.
Veröffentlicht: (2026)
von: Deichler, Anna, et al.
Veröffentlicht: (2026)
Evaluation Before Generation: A Paradigm for Robust Multimodal Sentiment Analysis with Missing Modalities
von: Chen, Rongfei, et al.
Veröffentlicht: (2026)
von: Chen, Rongfei, et al.
Veröffentlicht: (2026)
Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review
von: Rombach, Alexander Michael, et al.
Veröffentlicht: (2024)
von: Rombach, Alexander Michael, et al.
Veröffentlicht: (2024)
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
von: Li, Jianing, et al.
Veröffentlicht: (2024)
von: Li, Jianing, et al.
Veröffentlicht: (2024)
Depthwise Separable Convolutions with Deep Residual Convolutions
von: Hasan, Md Arid, et al.
Veröffentlicht: (2024)
von: Hasan, Md Arid, et al.
Veröffentlicht: (2024)
Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis
von: Heyne, Catyana, et al.
Veröffentlicht: (2026)
von: Heyne, Catyana, et al.
Veröffentlicht: (2026)
GroundCap: A Visually Grounded Image Captioning Dataset
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
von: Dehghani, Mahshid, et al.
Veröffentlicht: (2024)
von: Dehghani, Mahshid, et al.
Veröffentlicht: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Self-Supervised Borrowing Detection on Multilingual Wordlists
von: Wientzek, Tim
Veröffentlicht: (2025)
von: Wientzek, Tim
Veröffentlicht: (2025)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
von: He, Wei
Veröffentlicht: (2026)
von: He, Wei
Veröffentlicht: (2026)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
von: Freitas, Diogo, et al.
Veröffentlicht: (2025)
von: Freitas, Diogo, et al.
Veröffentlicht: (2025)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
von: Shah, Nisarg A., et al.
Veröffentlicht: (2025)
von: Shah, Nisarg A., et al.
Veröffentlicht: (2025)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
von: Cao, Jingtao, et al.
Veröffentlicht: (2024)
von: Cao, Jingtao, et al.
Veröffentlicht: (2024)
FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
CCS: Clinical Consensus Selection for Radiology Report Generation
von: Zhang, Xi, et al.
Veröffentlicht: (2026)
von: Zhang, Xi, et al.
Veröffentlicht: (2026)
GAEA: A Geolocation Aware Conversational Assistant
von: Campos, Ron, et al.
Veröffentlicht: (2025)
von: Campos, Ron, et al.
Veröffentlicht: (2025)
Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals
von: Baghel, Shruti Singh, et al.
Veröffentlicht: (2025)
von: Baghel, Shruti Singh, et al.
Veröffentlicht: (2025)
The American Sign Language Knowledge Graph: Infusing ASL Models with Linguistic Knowledge
von: Kezar, Lee, et al.
Veröffentlicht: (2024)
von: Kezar, Lee, et al.
Veröffentlicht: (2024)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
von: Cui, Shaoyang, et al.
Veröffentlicht: (2026)
von: Cui, Shaoyang, et al.
Veröffentlicht: (2026)
Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent
von: Yasuno, Takato
Veröffentlicht: (2026)
von: Yasuno, Takato
Veröffentlicht: (2026)
UVDoc: Neural Grid-based Document Unwarping
von: Verhoeven, Floor, et al.
Veröffentlicht: (2023)
von: Verhoeven, Floor, et al.
Veröffentlicht: (2023)
Leveraging Large Language Models for Semantic Query Processing in a Scholarly Knowledge Graph
von: Jia, Runsong, et al.
Veröffentlicht: (2024)
von: Jia, Runsong, et al.
Veröffentlicht: (2024)
Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content
von: Koushik, Girish A., et al.
Veröffentlicht: (2025)
von: Koushik, Girish A., et al.
Veröffentlicht: (2025)
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
von: Parcalabescu, Letitia, et al.
Veröffentlicht: (2021)
von: Parcalabescu, Letitia, et al.
Veröffentlicht: (2021)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
PlotPick: AI-powered batch extraction of numerical data from scientific figures
von: Carstensen, Tommy
Veröffentlicht: (2026)
von: Carstensen, Tommy
Veröffentlicht: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation
von: Yang, Qi, et al.
Veröffentlicht: (2026) -
Enhancing OCR for Sino-Vietnamese Language Processing via Fine-tuned PaddleOCRv5
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025) -
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2025) -
Large Language Models for Simultaneous Named Entity Extraction and Spelling Correction
von: Whittaker, Edward, et al.
Veröffentlicht: (2024) -
A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports
von: Schäfer, Henning, et al.
Veröffentlicht: (2025)