Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Maryam, Hiba, Fu, Ling, Song, Jiajun, Shafayet, Tajrian ABM, Luo, Qidi, Bai, Xiang, Liu, Yuliang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The First Swahili Language Scene Text Detection and Recognition Dataset
by: Douamba, Fadila Wendigoundi, et al.
Published: (2024)
by: Douamba, Fadila Wendigoundi, et al.
Published: (2024)
MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
by: Tang, Jingqun, et al.
Published: (2024)
by: Tang, Jingqun, et al.
Published: (2024)
UQA: Corpus for Urdu Question Answering
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
by: Al-Mohannadi, Aisha, et al.
Published: (2026)
by: Al-Mohannadi, Aisha, et al.
Published: (2026)
Ax-to-Grind Urdu: Benchmark Dataset for Urdu Fake News Detection
by: Harris, Sheetal, et al.
Published: (2024)
by: Harris, Sheetal, et al.
Published: (2024)
Efficient Visual Question Answering Pipeline for Autonomous Driving via Scene Region Compression
by: Cai, Yuliang, et al.
Published: (2026)
by: Cai, Yuliang, et al.
Published: (2026)
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
by: Shen, Zhixuan, et al.
Published: (2024)
by: Shen, Zhixuan, et al.
Published: (2024)
LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
by: Faisal, Faizan, et al.
Published: (2024)
by: Faisal, Faizan, et al.
Published: (2024)
Detection of Anxiety Levels in Urdu Text (Multiclass Dataset)
by: Fareed, Sadia, et al.
Published: (2025)
by: Fareed, Sadia, et al.
Published: (2025)
Scene-Text Grounding for Text-Based Video Question Answering
by: Zhou, Sheng, et al.
Published: (2024)
by: Zhou, Sheng, et al.
Published: (2024)
BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition
by: Ali, Wazir, et al.
Published: (2026)
by: Ali, Wazir, et al.
Published: (2026)
A Benchmark Dataset and a Framework for Urdu Multimodal Named Entity Recognition
by: Ahmad, Hussain, et al.
Published: (2025)
by: Ahmad, Hussain, et al.
Published: (2025)
Toward Real Text Manipulation Detection: New Dataset and New Solution
by: Luo, Dongliang, et al.
Published: (2023)
by: Luo, Dongliang, et al.
Published: (2023)
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
by: Fu, Ling, et al.
Published: (2024)
by: Fu, Ling, et al.
Published: (2024)
CluMo: Cluster-based Modality Fusion Prompt for Continual Learning in Visual Question Answering
by: Cai, Yuliang, et al.
Published: (2024)
by: Cai, Yuliang, et al.
Published: (2024)
STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes
by: Ishihara, Keishi, et al.
Published: (2025)
by: Ishihara, Keishi, et al.
Published: (2025)
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
by: Pham, Huy Quang, et al.
Published: (2024)
by: Pham, Huy Quang, et al.
Published: (2024)
MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
by: Shan, Bin, et al.
Published: (2024)
by: Shan, Bin, et al.
Published: (2024)
Exploration of Deep Learning Based Recognition for Urdu Text
by: Fazal, Sumaiya, et al.
Published: (2025)
by: Fazal, Sumaiya, et al.
Published: (2025)
From Press to Pixels: Evolving Urdu Text Recognition
by: Arif, Samee, et al.
Published: (2025)
by: Arif, Samee, et al.
Published: (2025)
How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking
by: Ahmed, Rafid, et al.
Published: (2026)
by: Ahmed, Rafid, et al.
Published: (2026)
3D Question Answering for City Scene Understanding
by: Sun, Penglei, et al.
Published: (2024)
by: Sun, Penglei, et al.
Published: (2024)
ZeShot-VQA: Zero-Shot Visual Question Answering Framework with Answer Mapping for Natural Disaster Damage Assessment
by: Karimi, Ehsan, et al.
Published: (2025)
by: Karimi, Ehsan, et al.
Published: (2025)
Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
by: Fu, Zijian, et al.
Published: (2025)
by: Fu, Zijian, et al.
Published: (2025)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
by: Yin, Liang, et al.
Published: (2025)
by: Yin, Liang, et al.
Published: (2025)
Evaluating Variance in Visual Question Answering Benchmarks
by: SR, Nikitha
Published: (2025)
by: SR, Nikitha
Published: (2025)
Hallucination Benchmark in Medical Visual Question Answering
by: Wu, Jinge, et al.
Published: (2024)
by: Wu, Jinge, et al.
Published: (2024)
Progressive Evolution from Single-Point to Polygon for Scene Text
by: Deng, Linger, et al.
Published: (2023)
by: Deng, Linger, et al.
Published: (2023)
MULTITAT: Benchmarking Multilingual Table-and-Text Question Answering
by: Zhang, Xuanliang, et al.
Published: (2025)
by: Zhang, Xuanliang, et al.
Published: (2025)
A Benchmark Dataset with Larger Context for Non-Factoid Question Answering over Islamic Text
by: Qamar, Faiza, et al.
Published: (2024)
by: Qamar, Faiza, et al.
Published: (2024)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
by: Qian, Tianwen, et al.
Published: (2023)
by: Qian, Tianwen, et al.
Published: (2023)
Text-to-TrajVis: Enabling Trajectory Data Visualizations from Natural Language Questions
by: Bai, Tian, et al.
Published: (2025)
by: Bai, Tian, et al.
Published: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
by: Nguyen, Hieu Minh, et al.
Published: (2025)
by: Nguyen, Hieu Minh, et al.
Published: (2025)
DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning
by: Yilmaz, Abdurrahim, et al.
Published: (2026)
by: Yilmaz, Abdurrahim, et al.
Published: (2026)
OWLViz: An Open-World Benchmark for Visual Question Answering
by: Nguyen, Thuy, et al.
Published: (2025)
by: Nguyen, Thuy, et al.
Published: (2025)
IPQA: A Benchmark for Core Intent Identification in Personalized Question Answering
by: Kim, Jieyong, et al.
Published: (2025)
by: Kim, Jieyong, et al.
Published: (2025)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
by: Song, Seokwon, et al.
Published: (2025)
by: Song, Seokwon, et al.
Published: (2025)
Similar Items
-
The First Swahili Language Scene Text Detection and Recognition Dataset
by: Douamba, Fadila Wendigoundi, et al.
Published: (2024) -
MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
by: Tang, Jingqun, et al.
Published: (2024) -
UQA: Corpus for Urdu Question Answering
by: Arif, Samee, et al.
Published: (2024) -
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
by: Al-Mohannadi, Aisha, et al.
Published: (2026) -
Ax-to-Grind Urdu: Benchmark Dataset for Urdu Fake News Detection
by: Harris, Sheetal, et al.
Published: (2024)