IndicSTR12: A Dataset for Indic Scene Text Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lunia, Harsh, Mondal, Ajoy, Jawahar, C V |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Deployable OCR models for Indic languages
von: Mathew, Minesh, et al.
Veröffentlicht: (2022)
von: Mathew, Minesh, et al.
Veröffentlicht: (2022)
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
von: Pal, Aniket, et al.
Veröffentlicht: (2024)
von: Pal, Aniket, et al.
Veröffentlicht: (2024)
PLATTER: A Page-Level Handwritten Text Recognition System for Indic Scripts
von: Kasuba, Badri Vishal, et al.
Veröffentlicht: (2025)
von: Kasuba, Badri Vishal, et al.
Veröffentlicht: (2025)
Can VLMs be used on videos for action recognition? LLMs are Visual Reasoning Coordinators
von: Lunia, Harsh
Veröffentlicht: (2024)
von: Lunia, Harsh
Veröffentlicht: (2024)
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
von: Pal, Aniket, et al.
Veröffentlicht: (2025)
von: Pal, Aniket, et al.
Veröffentlicht: (2025)
Navigating Text-to-Image Generative Bias across Indic Languages
von: Mittal, Surbhi, et al.
Veröffentlicht: (2024)
von: Mittal, Surbhi, et al.
Veröffentlicht: (2024)
MDiff4STR: Mask Diffusion Model for Scene Text Recognition
von: Du, Yongkun, et al.
Veröffentlicht: (2025)
von: Du, Yongkun, et al.
Veröffentlicht: (2025)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
von: Maniyar, Suyash, et al.
Veröffentlicht: (2025)
von: Maniyar, Suyash, et al.
Veröffentlicht: (2025)
DiffSTR: Controlled Diffusion Models for Scene Text Removal
von: Pathak, Sanhita, et al.
Veröffentlicht: (2024)
von: Pathak, Sanhita, et al.
Veröffentlicht: (2024)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
IndicFairFace: Balanced Indian Face Dataset for Auditing and Mitigating Geographical Bias in Vision-Language Models
von: Mohsin, Aarish Shah, et al.
Veröffentlicht: (2026)
von: Mohsin, Aarish Shah, et al.
Veröffentlicht: (2026)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
von: Faraz, Ali, et al.
Veröffentlicht: (2025)
von: Faraz, Ali, et al.
Veröffentlicht: (2025)
The First Swahili Language Scene Text Detection and Recognition Dataset
von: Douamba, Fadila Wendigoundi, et al.
Veröffentlicht: (2024)
von: Douamba, Fadila Wendigoundi, et al.
Veröffentlicht: (2024)
GaitSTR: Gait Recognition with Sequential Two-stream Refinement
von: Zheng, Wanrong, et al.
Veröffentlicht: (2024)
von: Zheng, Wanrong, et al.
Veröffentlicht: (2024)
STR-Cert: Robustness Certification for Deep Text Recognition on Deep Learning Pipelines and Vision Transformers
von: Shao, Daqian, et al.
Veröffentlicht: (2023)
von: Shao, Daqian, et al.
Veröffentlicht: (2023)
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
von: Maryam, Hiba, et al.
Veröffentlicht: (2024)
von: Maryam, Hiba, et al.
Veröffentlicht: (2024)
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
Reading Between the Lanes: Text VideoQA on the Road
von: Tom, George, et al.
Veröffentlicht: (2023)
von: Tom, George, et al.
Veröffentlicht: (2023)
Instruction-Guided Scene Text Recognition
von: Du, Yongkun, et al.
Veröffentlicht: (2024)
von: Du, Yongkun, et al.
Veröffentlicht: (2024)
Recognition-Synergistic Scene Text Editing
von: Fang, Zhengyao, et al.
Veröffentlicht: (2025)
von: Fang, Zhengyao, et al.
Veröffentlicht: (2025)
Decoder Pre-Training with only Text for Scene Text Recognition
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
TEACH: Text Encoding as Curriculum Hints for Scene Text Recognition
von: Yang, Xiahan, et al.
Veröffentlicht: (2025)
von: Yang, Xiahan, et al.
Veröffentlicht: (2025)
TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition
von: Ye, Xingsong, et al.
Veröffentlicht: (2024)
von: Ye, Xingsong, et al.
Veröffentlicht: (2024)
Reading in the Dark: Low-light Scene Text Recognition
von: Fu, Xuanshuo, et al.
Veröffentlicht: (2026)
von: Fu, Xuanshuo, et al.
Veröffentlicht: (2026)
Efficient and Accurate Scene Text Recognition with Cascaded-Transformers
von: Ozkan, Savas, et al.
Veröffentlicht: (2025)
von: Ozkan, Savas, et al.
Veröffentlicht: (2025)
Attend to what I say: Highlighting relevant content on slides
von: M, Megha Mariam K, et al.
Veröffentlicht: (2026)
von: M, Megha Mariam K, et al.
Veröffentlicht: (2026)
Source-free Video Domain Adaptation by Learning from Noisy Labels
von: Dasgupta, Avijit, et al.
Veröffentlicht: (2023)
von: Dasgupta, Avijit, et al.
Veröffentlicht: (2023)
MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition
von: Kassab, Hozaifa, et al.
Veröffentlicht: (2024)
von: Kassab, Hozaifa, et al.
Veröffentlicht: (2024)
A Dataset for Semantic Segmentation in the Presence of Unknowns
von: Laskar, Zakaria, et al.
Veröffentlicht: (2025)
von: Laskar, Zakaria, et al.
Veröffentlicht: (2025)
Surveying Facial Recognition Models for Diverse Indian Demographics: A Comparative Analysis on LFW and Custom Dataset
von: Pant, Pranav, et al.
Veröffentlicht: (2024)
von: Pant, Pranav, et al.
Veröffentlicht: (2024)
Masked and Permuted Implicit Context Learning for Scene Text Recognition
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2023)
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2023)
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
von: Jahagirdar, Soumya Shamarao, et al.
Veröffentlicht: (2025)
von: Jahagirdar, Soumya Shamarao, et al.
Veröffentlicht: (2025)
Real-Time Human Reconstruction and Animation using Feed-Forward Gaussian Splatting
von: Chatterjee, Devdoot, et al.
Veröffentlicht: (2026)
von: Chatterjee, Devdoot, et al.
Veröffentlicht: (2026)
CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition
von: Zheng, Jinzhi, et al.
Veröffentlicht: (2024)
von: Zheng, Jinzhi, et al.
Veröffentlicht: (2024)
SVIPTR: Fast and Efficient Scene Text Recognition with Vision Permutable Extractor
von: Cheng, Xianfu, et al.
Veröffentlicht: (2024)
von: Cheng, Xianfu, et al.
Veröffentlicht: (2024)
Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition
von: Lin, Tiancheng, et al.
Veröffentlicht: (2024)
von: Lin, Tiancheng, et al.
Veröffentlicht: (2024)
Boosting Semi-Supervised Scene Text Recognition via Viewing and Summarizing
von: Qu, Yadong, et al.
Veröffentlicht: (2024)
von: Qu, Yadong, et al.
Veröffentlicht: (2024)
Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition
von: Zhou, Bangbang, et al.
Veröffentlicht: (2024)
von: Zhou, Bangbang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Deployable OCR models for Indic languages
von: Mathew, Minesh, et al.
Veröffentlicht: (2022) -
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
von: Pal, Aniket, et al.
Veröffentlicht: (2024) -
PLATTER: A Page-Level Handwritten Text Recognition System for Indic Scripts
von: Kasuba, Badri Vishal, et al.
Veröffentlicht: (2025) -
Can VLMs be used on videos for action recognition? LLMs are Visual Reasoning Coordinators
von: Lunia, Harsh
Veröffentlicht: (2024) -
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
von: Pal, Aniket, et al.
Veröffentlicht: (2025)