COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
Fuente:
arXiv
Salvato in:
| Autore principale: | Hassan, Umair |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
di: Shin, Suyeon, et al.
Pubblicazione: (2024)
di: Shin, Suyeon, et al.
Pubblicazione: (2024)
If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions
di: Barbano, Carlo Alberto, et al.
Pubblicazione: (2024)
di: Barbano, Carlo Alberto, et al.
Pubblicazione: (2024)
BAAF: Universal Transformation of One-Class Classifiers for Unsupervised Image Anomaly Detection
di: McIntosh, Declan, et al.
Pubblicazione: (2026)
di: McIntosh, Declan, et al.
Pubblicazione: (2026)
VLM@school -- Evaluation of AI image understanding on German middle school knowledge
di: Peinl, René, et al.
Pubblicazione: (2025)
di: Peinl, René, et al.
Pubblicazione: (2025)
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
di: Brusnicki, Roberto, et al.
Pubblicazione: (2026)
di: Brusnicki, Roberto, et al.
Pubblicazione: (2026)
Poisson Flow Consistency Training
di: Zhang, Anthony, et al.
Pubblicazione: (2025)
di: Zhang, Anthony, et al.
Pubblicazione: (2025)
CNNtention: Can CNNs do better with Attention?
di: Kapila, Nikhil, et al.
Pubblicazione: (2024)
di: Kapila, Nikhil, et al.
Pubblicazione: (2024)
Does CLIP perceive art the same way we do?
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
di: Faisal, Faizan, et al.
Pubblicazione: (2024)
di: Faisal, Faizan, et al.
Pubblicazione: (2024)
Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
di: Gkountouras, John, et al.
Pubblicazione: (2025)
di: Gkountouras, John, et al.
Pubblicazione: (2025)
Application of deep learning approaches for medieval historical documents transcription
di: Voloshchuk, Maksym, et al.
Pubblicazione: (2025)
di: Voloshchuk, Maksym, et al.
Pubblicazione: (2025)
Search Multilayer Perceptron-Based Fusion for Efficient and Accurate Siamese Tracking
di: Shen, Tianqi, et al.
Pubblicazione: (2026)
di: Shen, Tianqi, et al.
Pubblicazione: (2026)
MambaNetLK: Enhancing Colonoscopy Point Cloud Registration with Mamba
di: Jiang, Linzhe, et al.
Pubblicazione: (2025)
di: Jiang, Linzhe, et al.
Pubblicazione: (2025)
LRVS-Fashion: Extending Visual Search with Referring Instructions
di: Lepage, Simon, et al.
Pubblicazione: (2023)
di: Lepage, Simon, et al.
Pubblicazione: (2023)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
di: Yanambakkam, Hemanth Teja, et al.
Pubblicazione: (2025)
di: Yanambakkam, Hemanth Teja, et al.
Pubblicazione: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
di: Ma, Yueen, et al.
Pubblicazione: (2024)
di: Ma, Yueen, et al.
Pubblicazione: (2024)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
di: Li, Mingda, et al.
Pubblicazione: (2024)
di: Li, Mingda, et al.
Pubblicazione: (2024)
Multimodal Image Matching based on Frequency-domain Information of Local Energy Response
di: Yang, Meng, et al.
Pubblicazione: (2025)
di: Yang, Meng, et al.
Pubblicazione: (2025)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
di: Dumpala, Sri Harsha, et al.
Pubblicazione: (2024)
di: Dumpala, Sri Harsha, et al.
Pubblicazione: (2024)
LayerAct: Advanced Activation Mechanism for Robust Inference of CNNs
di: Yoon, Kihyuk, et al.
Pubblicazione: (2023)
di: Yoon, Kihyuk, et al.
Pubblicazione: (2023)
HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models
di: Zhang, Shengkai, et al.
Pubblicazione: (2024)
di: Zhang, Shengkai, et al.
Pubblicazione: (2024)
Overcoming Catastrophic Forgetting in Federated Class-Incremental Learning via Federated Global Twin Generator
di: Nguyen, Thinh, et al.
Pubblicazione: (2024)
di: Nguyen, Thinh, et al.
Pubblicazione: (2024)
Revisiting Sampson Approximations for Geometric Estimation Problems
di: Rydell, Felix, et al.
Pubblicazione: (2024)
di: Rydell, Felix, et al.
Pubblicazione: (2024)
Nearest Neighbor Projection Removal Adversarial Training
di: Singh, Himanshu, et al.
Pubblicazione: (2025)
di: Singh, Himanshu, et al.
Pubblicazione: (2025)
Unsupervised Anomaly Detection Using Diffusion Trend Analysis for Display Inspection
di: Kim, Eunwoo, et al.
Pubblicazione: (2024)
di: Kim, Eunwoo, et al.
Pubblicazione: (2024)
ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
di: Liu, Xueyi, et al.
Pubblicazione: (2025)
di: Liu, Xueyi, et al.
Pubblicazione: (2025)
Empowering Image Recovery_ A Multi-Attention Approach
di: Wen, Juan, et al.
Pubblicazione: (2024)
di: Wen, Juan, et al.
Pubblicazione: (2024)
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
di: Malikussaid, et al.
Pubblicazione: (2026)
di: Malikussaid, et al.
Pubblicazione: (2026)
Benchmarking Vision Language Models on German Factual Data
di: Peinl, René, et al.
Pubblicazione: (2025)
di: Peinl, René, et al.
Pubblicazione: (2025)
Semi-automated extraction of research topics and trends from NCI funding in radiological sciences from 2000-2020
di: Nguyen, Mark, et al.
Pubblicazione: (2023)
di: Nguyen, Mark, et al.
Pubblicazione: (2023)
Archival Faces: Detection of Faces in Digitized Historical Documents
di: Vaško, Marek, et al.
Pubblicazione: (2025)
di: Vaško, Marek, et al.
Pubblicazione: (2025)
Generation of maximal snake polyominoes using a deep neural network
di: Gauthier, Benjamin, et al.
Pubblicazione: (2026)
di: Gauthier, Benjamin, et al.
Pubblicazione: (2026)
How many samples to label for an application given a foundation model? Chest X-ray classification study
di: Nechaev, Nikolay, et al.
Pubblicazione: (2025)
di: Nechaev, Nikolay, et al.
Pubblicazione: (2025)
JVLGS: Joint Vision-Language Gas Leak Segmentation
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)
VALE: A Multimodal Visual and Language Explanation Framework for Image Classifiers using eXplainable AI and Language Models
di: Natarajan, Purushothaman, et al.
Pubblicazione: (2024)
di: Natarajan, Purushothaman, et al.
Pubblicazione: (2024)
FastGS: Training 3D Gaussian Splatting in 100 Seconds
di: Ren, Shiwei, et al.
Pubblicazione: (2025)
di: Ren, Shiwei, et al.
Pubblicazione: (2025)
SEGS-SLAM: Structure-enhanced 3D Gaussian Splatting SLAM with Appearance Embedding
di: Wen, Tianci, et al.
Pubblicazione: (2025)
di: Wen, Tianci, et al.
Pubblicazione: (2025)
A large-scale, physically-based synthetic dataset for satellite pose estimation
di: Velkei, Szabolcs, et al.
Pubblicazione: (2025)
di: Velkei, Szabolcs, et al.
Pubblicazione: (2025)
QCFace: Image Quality Control for boosting Face Representation & Recognition
di: Doan-Ngo, Duc-Phuong, et al.
Pubblicazione: (2025)
di: Doan-Ngo, Duc-Phuong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
di: Shin, Suyeon, et al.
Pubblicazione: (2024) -
If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions
di: Barbano, Carlo Alberto, et al.
Pubblicazione: (2024) -
BAAF: Universal Transformation of One-Class Classifiers for Unsupervised Image Anomaly Detection
di: McIntosh, Declan, et al.
Pubblicazione: (2026) -
VLM@school -- Evaluation of AI image understanding on German middle school knowledge
di: Peinl, René, et al.
Pubblicazione: (2025) -
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
di: Brusnicki, Roberto, et al.
Pubblicazione: (2026)