LeafNet: A Large-Scale Dataset and Comprehensive Benchmark for Foundational Vision-Language Understanding of Plant Diseases
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Quoc, Khang Nguyen, Dao, Phuong D., Quach, Luyl-Da |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Vision-Language Foundation Model for Leaf Disease Identification
von: Quoc, Khang Nguyen, et al.
Veröffentlicht: (2025)
von: Quoc, Khang Nguyen, et al.
Veröffentlicht: (2025)
Maize Leaf Disease
von: Luyl-Da, Quach
Veröffentlicht: (2025)
von: Luyl-Da, Quach
Veröffentlicht: (2025)
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
von: Tran, Quoc-Khang, et al.
Veröffentlicht: (2026)
von: Tran, Quoc-Khang, et al.
Veröffentlicht: (2026)
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2025)
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2025)
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
von: Zhang, Da, et al.
Veröffentlicht: (2025)
von: Zhang, Da, et al.
Veröffentlicht: (2025)
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
von: Nguyen, Huu Tien, et al.
Veröffentlicht: (2025)
von: Nguyen, Huu Tien, et al.
Veröffentlicht: (2025)
PlantSeg: A Large-Scale In-the-wild Dataset for Plant Disease Segmentation
von: Wei, Tianqi, et al.
Veröffentlicht: (2024)
von: Wei, Tianqi, et al.
Veröffentlicht: (2024)
Leveraging Model Soups to Classify Intangible Cultural Heritage Images from the Mekong Delta
von: Tran, Quoc-Khang, et al.
Veröffentlicht: (2026)
von: Tran, Quoc-Khang, et al.
Veröffentlicht: (2026)
ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models
von: Yuan, Zhenghang, et al.
Veröffentlicht: (2024)
von: Yuan, Zhenghang, et al.
Veröffentlicht: (2024)
Insect-Foundation: A Foundation Model and Large-scale 1M Dataset for Visual Insect Understanding
von: Nguyen, Hoang-Quan, et al.
Veröffentlicht: (2023)
von: Nguyen, Hoang-Quan, et al.
Veröffentlicht: (2023)
IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
von: Pham, Huy Quang, et al.
Veröffentlicht: (2024)
von: Pham, Huy Quang, et al.
Veröffentlicht: (2024)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding
von: Hu, Ming, et al.
Veröffentlicht: (2024)
von: Hu, Ming, et al.
Veröffentlicht: (2024)
A Comprehensive Study on Medical Image Segmentation using Deep Neural Networks
von: Dao, Loan, et al.
Veröffentlicht: (2025)
von: Dao, Loan, et al.
Veröffentlicht: (2025)
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
Resampling Benchmark for Efficient Comprehensive Evaluation of Large Vision-Language Models
von: Suzuki, Teppei, et al.
Veröffentlicht: (2025)
von: Suzuki, Teppei, et al.
Veröffentlicht: (2025)
SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting
von: Ma, Mengjiao, et al.
Veröffentlicht: (2025)
von: Ma, Mengjiao, et al.
Veröffentlicht: (2025)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
LeafTrackNet: A Deep Learning Framework for Robust Leaf Tracking in Top-Down Plant Phenotyping
von: Liu, Shanghua, et al.
Veröffentlicht: (2025)
von: Liu, Shanghua, et al.
Veröffentlicht: (2025)
LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models
von: Qin, Zhenyue, et al.
Veröffentlicht: (2024)
von: Qin, Zhenyue, et al.
Veröffentlicht: (2024)
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2026)
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2026)
DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark
von: Li, Haodong, et al.
Veröffentlicht: (2024)
von: Li, Haodong, et al.
Veröffentlicht: (2024)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
von: Nguyen, Thong, et al.
Veröffentlicht: (2025)
von: Nguyen, Thong, et al.
Veröffentlicht: (2025)
FloraSyntropy-Net: Scalable Deep Learning with Novel FloraSyntropy Archive for Large-Scale Plant Disease Diagnosis
von: Khan, Saif Ur Rehman, et al.
Veröffentlicht: (2025)
von: Khan, Saif Ur Rehman, et al.
Veröffentlicht: (2025)
ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
PM25Vision: A Large-Scale Benchmark Dataset for Visual Estimation of Air Quality
von: Han, Yang
Veröffentlicht: (2025)
von: Han, Yang
Veröffentlicht: (2025)
ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving
von: Wang, Yujin, et al.
Veröffentlicht: (2026)
von: Wang, Yujin, et al.
Veröffentlicht: (2026)
U-Bench: A Comprehensive Understanding of U-Net through 100-Variant Benchmarking
von: Tang, Fenghe, et al.
Veröffentlicht: (2025)
von: Tang, Fenghe, et al.
Veröffentlicht: (2025)
ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
von: Felemban, Abdulwahab, et al.
Veröffentlicht: (2025)
von: Felemban, Abdulwahab, et al.
Veröffentlicht: (2025)
LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
von: Qin, Zhenyue, et al.
Veröffentlicht: (2025)
von: Qin, Zhenyue, et al.
Veröffentlicht: (2025)
PlantVillageVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
von: Sakib, Syed Nazmus, et al.
Veröffentlicht: (2025)
von: Sakib, Syed Nazmus, et al.
Veröffentlicht: (2025)
EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
von: Li, Sijing, et al.
Veröffentlicht: (2025)
von: Li, Sijing, et al.
Veröffentlicht: (2025)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
von: Jiang, Siyang, et al.
Veröffentlicht: (2025)
von: Jiang, Siyang, et al.
Veröffentlicht: (2025)
DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
von: Liu, Zixuan, et al.
Veröffentlicht: (2025)
von: Liu, Zixuan, et al.
Veröffentlicht: (2025)
Towards All-Day Perception for Off-Road Driving: A Large-Scale Multispectral Dataset and Comprehensive Benchmark
von: Wang, Shuo, et al.
Veröffentlicht: (2026)
von: Wang, Shuo, et al.
Veröffentlicht: (2026)
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2025)
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Vision-Language Foundation Model for Leaf Disease Identification
von: Quoc, Khang Nguyen, et al.
Veröffentlicht: (2025) -
Maize Leaf Disease
von: Luyl-Da, Quach
Veröffentlicht: (2025) -
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
von: Tran, Quoc-Khang, et al.
Veröffentlicht: (2026) -
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2025) -
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
von: Zhang, Da, et al.
Veröffentlicht: (2025)