What Lies Beneath: A Call for Distribution-based Visual Question & Answer Datasets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Naiman, Jill P., Evans, Daniel J., Seo, JooYoung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hybrid Retrieval-Augmented Generation for Robust Multilingual Document Question Answering
von: Mudet, Anthony, et al.
Veröffentlicht: (2025)
von: Mudet, Anthony, et al.
Veröffentlicht: (2025)
WildDepth: A Multimodal Dataset for 3D Wildlife Perception and Depth Estimation
von: Aamir, Muhammad, et al.
Veröffentlicht: (2026)
von: Aamir, Muhammad, et al.
Veröffentlicht: (2026)
Knowledge Graphs for Digitized Manuscripts in Jagiellonian Digital Library Application
von: Ignatowicz, Jan, et al.
Veröffentlicht: (2025)
von: Ignatowicz, Jan, et al.
Veröffentlicht: (2025)
Unfolding the Past: A Comprehensive Deep Learning Approach to Analyzing Incunabula Pages
von: Ropel, Klaudia, et al.
Veröffentlicht: (2025)
von: Ropel, Klaudia, et al.
Veröffentlicht: (2025)
Ink Detection from Surface Topography of the Herculaneum Papyri
von: Angelotti, Giorgio, et al.
Veröffentlicht: (2026)
von: Angelotti, Giorgio, et al.
Veröffentlicht: (2026)
Structured Analysis and Comparison of Alphabets in Historical Handwritten Ciphers
von: Méndez, Martín, et al.
Veröffentlicht: (2024)
von: Méndez, Martín, et al.
Veröffentlicht: (2024)
Automatic Recognition of Learning Resource Category in a Digital Library
von: Banerjee, Soumya, et al.
Veröffentlicht: (2023)
von: Banerjee, Soumya, et al.
Veröffentlicht: (2023)
LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR
von: Yang, Guang, et al.
Veröffentlicht: (2025)
von: Yang, Guang, et al.
Veröffentlicht: (2025)
An Intelligent Framework for Real-Time Yoga Pose Detection and Posture Correction
von: Haldar, Chandramouli
Veröffentlicht: (2026)
von: Haldar, Chandramouli
Veröffentlicht: (2026)
Productivity profile of CNPq scholarship researchers in computer science from 2017 to 2021
von: Albertini, Marcelo Keese, et al.
Veröffentlicht: (2024)
von: Albertini, Marcelo Keese, et al.
Veröffentlicht: (2024)
μgat: Improving Single-Page Document Parsing by Providing Multi-Page Context
von: Quattrini, Fabio, et al.
Veröffentlicht: (2024)
von: Quattrini, Fabio, et al.
Veröffentlicht: (2024)
NeuroPapyri: A Deep Attention Embedding Network for Handwritten Papyri Retrieval
von: De Gregorio, Giuseppe, et al.
Veröffentlicht: (2024)
von: De Gregorio, Giuseppe, et al.
Veröffentlicht: (2024)
Quo Vadis Handwritten Text Generation for Handwritten Text Recognition?
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
Evolving Thematic Map Design in Academic Cartography: A Thirty-Year Study Based on Multilingual Journals
von: Wei, Zhiwei, et al.
Veröffentlicht: (2026)
von: Wei, Zhiwei, et al.
Veröffentlicht: (2026)
Segmentation of Ink and Parchment in Dead Sea Scroll Fragments
von: Kurar-Barakat, Berat, et al.
Veröffentlicht: (2024)
von: Kurar-Barakat, Berat, et al.
Veröffentlicht: (2024)
Callico: a Versatile Open-Source Document Image Annotation Platform
von: Kermorvant, Christopher, et al.
Veröffentlicht: (2024)
von: Kermorvant, Christopher, et al.
Veröffentlicht: (2024)
Automatic Reviewers Assignment to a Research Paper Based on Allied References and Publications Weight
von: Mahmud, Tamim Al, et al.
Veröffentlicht: (2025)
von: Mahmud, Tamim Al, et al.
Veröffentlicht: (2025)
A Survey of OCR Evaluation Methods and Metrics and the Invisibility of Historical Documents
von: Beyene, Fitsum Sileshi, et al.
Veröffentlicht: (2026)
von: Beyene, Fitsum Sileshi, et al.
Veröffentlicht: (2026)
[Citation needed] Data usage and citation practices in medical imaging conferences
von: Sourget, Théo, et al.
Veröffentlicht: (2024)
von: Sourget, Théo, et al.
Veröffentlicht: (2024)
Object Recognition from Scientific Document based on Compartment Refinement Framework
von: Li, Jinghong, et al.
Veröffentlicht: (2023)
von: Li, Jinghong, et al.
Veröffentlicht: (2023)
Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)
von: Griesshaber, Niclas, et al.
Veröffentlicht: (2025)
von: Griesshaber, Niclas, et al.
Veröffentlicht: (2025)
LEMUR Neural Network Dataset: Towards Seamless AutoML
von: Goodarzi, Arash Torabi, et al.
Veröffentlicht: (2025)
von: Goodarzi, Arash Torabi, et al.
Veröffentlicht: (2025)
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2026)
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2026)
Optical Music Recognition in Manuscripts from the Ricordi Archive
von: Simonetta, Federico, et al.
Veröffentlicht: (2024)
von: Simonetta, Federico, et al.
Veröffentlicht: (2024)
DeepScribe: Localization and Classification of Elamite Cuneiform Signs Via Deep Learning
von: Williams, Edward C., et al.
Veröffentlicht: (2023)
von: Williams, Edward C., et al.
Veröffentlicht: (2023)
A Global Atlas of Digital Dermatology to Map Innovation and Disparities
von: Gröger, Fabian, et al.
Veröffentlicht: (2025)
von: Gröger, Fabian, et al.
Veröffentlicht: (2025)
Sheet Music Benchmark: Standardized Optical Music Recognition Evaluation
von: Martinez-Sevilla, Juan C., et al.
Veröffentlicht: (2025)
von: Martinez-Sevilla, Juan C., et al.
Veröffentlicht: (2025)
Transfer Learning Approach for Railway Technical Map (RTM) Component Identification
von: Rumalshan, Obadage Rochana, et al.
Veröffentlicht: (2024)
von: Rumalshan, Obadage Rochana, et al.
Veröffentlicht: (2024)
An HTR-LLM Workflow for High-Accuracy Transcription and Analysis of Abbreviated Latin Court Hand
von: Isom, Joshua D.
Veröffentlicht: (2025)
von: Isom, Joshua D.
Veröffentlicht: (2025)
[Re] Network Deconvolution
von: Obadage, Rochana R., et al.
Veröffentlicht: (2024)
von: Obadage, Rochana R., et al.
Veröffentlicht: (2024)
A Literature Review of Literature Reviews in Pattern Analysis and Machine Intelligence
von: Zhao, Penghai, et al.
Veröffentlicht: (2024)
von: Zhao, Penghai, et al.
Veröffentlicht: (2024)
Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
von: Duan, Changxu
Veröffentlicht: (2025)
von: Duan, Changxu
Veröffentlicht: (2025)
Efficient OCR for Building a Diverse Digital History
von: Carlson, Jacob, et al.
Veröffentlicht: (2023)
von: Carlson, Jacob, et al.
Veröffentlicht: (2023)
Video Compression for Spatiotemporal Earth System Data
von: Pellicer-Valero, Oscar J., et al.
Veröffentlicht: (2025)
von: Pellicer-Valero, Oscar J., et al.
Veröffentlicht: (2025)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
In the Picture: Medical Imaging Datasets, Artifacts, and their Living Review
von: Jiménez-Sánchez, Amelia, et al.
Veröffentlicht: (2025)
von: Jiménez-Sánchez, Amelia, et al.
Veröffentlicht: (2025)
Sustaining Knowledge Infrastructures: Asking the Right Questions and Listening for Answers
von: Gregory, Kathleen, et al.
Veröffentlicht: (2025)
von: Gregory, Kathleen, et al.
Veröffentlicht: (2025)
Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
von: Wu, Yu, et al.
Veröffentlicht: (2025)
von: Wu, Yu, et al.
Veröffentlicht: (2025)
PubMed-OCR: PMC Open Access OCR Annotations
von: Heidenreich, Hunter, et al.
Veröffentlicht: (2026)
von: Heidenreich, Hunter, et al.
Veröffentlicht: (2026)
TRIDIS: A Comprehensive Medieval and Early Modern Corpus for HTR and NER
von: Aguilar, Sergio Torres
Veröffentlicht: (2025)
von: Aguilar, Sergio Torres
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hybrid Retrieval-Augmented Generation for Robust Multilingual Document Question Answering
von: Mudet, Anthony, et al.
Veröffentlicht: (2025) -
WildDepth: A Multimodal Dataset for 3D Wildlife Perception and Depth Estimation
von: Aamir, Muhammad, et al.
Veröffentlicht: (2026) -
Knowledge Graphs for Digitized Manuscripts in Jagiellonian Digital Library Application
von: Ignatowicz, Jan, et al.
Veröffentlicht: (2025) -
Unfolding the Past: A Comprehensive Deep Learning Approach to Analyzing Incunabula Pages
von: Ropel, Klaudia, et al.
Veröffentlicht: (2025) -
Ink Detection from Surface Topography of the Herculaneum Papyri
von: Angelotti, Giorgio, et al.
Veröffentlicht: (2026)