600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script
Fuente:
arXiv
Saved in:
| Main Author: | Malik, Haq Nawaz |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
by: Malik, Haq Nawaz, et al.
Published: (2026)
by: Malik, Haq Nawaz, et al.
Published: (2026)
ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining
by: Malik, Haq Nawaz
Published: (2026)
by: Malik, Haq Nawaz
Published: (2026)
ks-pret-5m: a 5 million word, 12 million token kashmiri pretraining dataset
by: Malik, Haq Nawaz, et al.
Published: (2026)
by: Malik, Haq Nawaz, et al.
Published: (2026)
Comparative analysis of optical character recognition methods for Sámi texts from the National Library of Norway
by: Enstad, Tita, et al.
Published: (2025)
by: Enstad, Tita, et al.
Published: (2025)
An open dataset for oracle bone script recognition and decipherment
by: Wang, Pengjie, et al.
Published: (2024)
by: Wang, Pengjie, et al.
Published: (2024)
A large-scale image-text dataset benchmark for farmland segmentation
by: Tao, Chao, et al.
Published: (2025)
by: Tao, Chao, et al.
Published: (2025)
A large-scale dataset for end-to-end table recognition in the wild
by: Yang, Fan, et al.
Published: (2023)
by: Yang, Fan, et al.
Published: (2023)
HCR-Net: A deep learning based script independent handwritten character recognition network
by: Chauhan, Vinod Kumar, et al.
Published: (2021)
by: Chauhan, Vinod Kumar, et al.
Published: (2021)
dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
by: Li, Yumeng, et al.
Published: (2025)
by: Li, Yumeng, et al.
Published: (2025)
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
by: Zhao, Tiancheng, et al.
Published: (2022)
by: Zhao, Tiancheng, et al.
Published: (2022)
FSboard: Over 3 million characters of ASL fingerspelling collected via smartphones
by: Georg, Manfred, et al.
Published: (2024)
by: Georg, Manfred, et al.
Published: (2024)
Learning based Ge'ez character handwritten recognition
by: Yimer, Hailemicael Lulseged, et al.
Published: (2024)
by: Yimer, Hailemicael Lulseged, et al.
Published: (2024)
Sumotosima: A Framework and Dataset for Classifying and Summarizing Otoscopic Images
by: Khan, Eram Anwarul, et al.
Published: (2024)
by: Khan, Eram Anwarul, et al.
Published: (2024)
Link prediction Graph Neural Networks for structure recognition of Handwritten Mathematical Expressions
by: Nguyen, Cuong Tuan, et al.
Published: (2025)
by: Nguyen, Cuong Tuan, et al.
Published: (2025)
CENSUS-HWR: a large training dataset for offline handwriting recognition
by: Joshi, Chetan, et al.
Published: (2023)
by: Joshi, Chetan, et al.
Published: (2023)
Approaches of large-scale images recognition with more than 50,000 categoris
by: Huang, Wanhong, et al.
Published: (2020)
by: Huang, Wanhong, et al.
Published: (2020)
Evaluating the plausibility of synthetic images for improving automated endoscopic stone recognition
by: Gonzalez-Perez, Ruben, et al.
Published: (2024)
by: Gonzalez-Perez, Ruben, et al.
Published: (2024)
Improving the generalization of gait recognition with limited datasets
by: Zhou, Qian, et al.
Published: (2025)
by: Zhou, Qian, et al.
Published: (2025)
HABD: a houma alliance book ancient handwritten character recognition database
by: Yuan, Xiaoyu, et al.
Published: (2024)
by: Yuan, Xiaoyu, et al.
Published: (2024)
An evaluation of Deep Learning based stereo dense matching dataset shift from aerial images and a large scale stereo dataset
by: Wu, Teng, et al.
Published: (2024)
by: Wu, Teng, et al.
Published: (2024)
KOLOMVERSE: Korea open large-scale image dataset for object detection in the maritime universe
by: Nanda, Abhilasha, et al.
Published: (2022)
by: Nanda, Abhilasha, et al.
Published: (2022)
rareboost3d: a synthetic lidar dataset with enhanced rare classes
by: Lin, Shutong, et al.
Published: (2025)
by: Lin, Shutong, et al.
Published: (2025)
GBSS:a global building semantic segmentation dataset for large-scale remote sensing building extraction
by: Hu, Yuping, et al.
Published: (2024)
by: Hu, Yuping, et al.
Published: (2024)
SYNBUILD-3D: A large, multi-modal, and semantically rich synthetic dataset of 3D building models at Level of Detail 4
by: Mayer, Kevin, et al.
Published: (2025)
by: Mayer, Kevin, et al.
Published: (2025)
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
by: Futeral, Matthieu, et al.
Published: (2024)
by: Futeral, Matthieu, et al.
Published: (2024)
A multimodal gesture recognition dataset for desktop human-computer interaction
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
A comprehensive survey of oracle character recognition: challenges, benchmarks, and beyond
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination
by: Zhong, Xinliu, et al.
Published: (2025)
by: Zhong, Xinliu, et al.
Published: (2025)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
by: Ashraf, Tajamul, et al.
Published: (2025)
by: Ashraf, Tajamul, et al.
Published: (2025)
Realism to Deception: Investigating Deepfake Detectors Against Face Enhancement
by: Saeed, Muhammad Saad, et al.
Published: (2025)
by: Saeed, Muhammad Saad, et al.
Published: (2025)
A Multitask Deep Learning Model for Classification and Regression of Hyperspectral Images: Application to the large-scale dataset
by: Chhapariya, Koushikey, et al.
Published: (2024)
by: Chhapariya, Koushikey, et al.
Published: (2024)
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Superhuman performance in urology board questions by an explainable large language model enabled for context integration of the European Association of Urology guidelines: the UroBot study
by: Hetz, Martin J., et al.
Published: (2024)
by: Hetz, Martin J., et al.
Published: (2024)
DocXPand-25k: a large and diverse benchmark dataset for identity documents analysis
by: Lerouge, Julien, et al.
Published: (2024)
by: Lerouge, Julien, et al.
Published: (2024)
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
by: Ashraf, Tajamul, et al.
Published: (2025)
by: Ashraf, Tajamul, et al.
Published: (2025)
RealSynCol: a high-fidelity synthetic colon dataset for 3D reconstruction applications
by: Lena, Chiara, et al.
Published: (2026)
by: Lena, Chiara, et al.
Published: (2026)
Learning from the few: Fine-grained approach to pediatric wrist pathology recognition on a limited dataset
by: Ahmed, Ammar, et al.
Published: (2024)
by: Ahmed, Ammar, et al.
Published: (2024)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
A workflow for generating synthetic LiDAR datasets in simulation environments
by: Phadke, Abhishek, et al.
Published: (2025)
by: Phadke, Abhishek, et al.
Published: (2025)
Three-dimensional visualization of X-ray micro-CT with large-scale datasets: Efficiency and accuracy for real-time interaction
by: Yin, Yipeng, et al.
Published: (2026)
by: Yin, Yipeng, et al.
Published: (2026)
Similar Items
-
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
by: Malik, Haq Nawaz, et al.
Published: (2026) -
ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining
by: Malik, Haq Nawaz
Published: (2026) -
ks-pret-5m: a 5 million word, 12 million token kashmiri pretraining dataset
by: Malik, Haq Nawaz, et al.
Published: (2026) -
Comparative analysis of optical character recognition methods for Sámi texts from the National Library of Norway
by: Enstad, Tita, et al.
Published: (2025) -
An open dataset for oracle bone script recognition and decipherment
by: Wang, Pengjie, et al.
Published: (2024)