A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiang, Liu, Zhaoxiang, Hu, Huan, Chen, Zezhou, Wang, Kohou, Wang, Kai, Lian, Shiguo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Large Vision-Language Model based Environment Perception System for Visually Impaired People
by: Chen, Zezhou, et al.
Published: (2025)
by: Chen, Zezhou, et al.
Published: (2025)
Hierarchical Deep Fusion Framework for Multi-dimensional Facial Forgery Detection -- The 2024 Global Deepfake Image Detection Challenge
by: Wang, Kohou, et al.
Published: (2025)
by: Wang, Kohou, et al.
Published: (2025)
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
by: Wang, Kohou, et al.
Published: (2025)
by: Wang, Kohou, et al.
Published: (2025)
MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
by: Zhao, Kaikai, et al.
Published: (2025)
by: Zhao, Kaikai, et al.
Published: (2025)
Patch-wise Auto-Encoder for Visual Anomaly Detection
by: Cui, Yajie, et al.
Published: (2023)
by: Cui, Yajie, et al.
Published: (2023)
PSTF-AttControl: Per-Subject-Tuning-Free Personalized Image Generation with Controllable Face Attributes
by: liu, Xiang, et al.
Published: (2025)
by: liu, Xiang, et al.
Published: (2025)
SLearnLLM: A Self-Learning Framework for Efficient Domain-Specific Adaptation of Large Language Models
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
by: Gao, Huanlin, et al.
Published: (2025)
by: Gao, Huanlin, et al.
Published: (2025)
Piculet: Specialized Models-Guided Hallucination Decrease for MultiModal Large Language Models
by: Wang, Kohou, et al.
Published: (2024)
by: Wang, Kohou, et al.
Published: (2024)
Chain-of-Trajectories: Unlocking the Intrinsic Generative Optimality of Diffusion Models via Graph-Theoretic Planning
by: Chen, Ping, et al.
Published: (2026)
by: Chen, Ping, et al.
Published: (2026)
Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3D
by: Chen, Ping, et al.
Published: (2026)
by: Chen, Ping, et al.
Published: (2026)
Sparser Block-Sparse Attention via Token Permutation
by: Wang, Xinghao, et al.
Published: (2025)
by: Wang, Xinghao, et al.
Published: (2025)
Optimizing for the Shortest Path in Denoising Diffusion Model
by: Chen, Ping, et al.
Published: (2025)
by: Chen, Ping, et al.
Published: (2025)
MM-NeuroOnco: A Multimodal Benchmark and Instruction Dataset for MRI-Based Brain Tumor Diagnosis
by: Guo, Feng, et al.
Published: (2026)
by: Guo, Feng, et al.
Published: (2026)
TP3M: Transformer-based Pseudo 3D Image Matching with Reference Image
by: Han, Liming, et al.
Published: (2024)
by: Han, Liming, et al.
Published: (2024)
KAConvNet: Kolmogorov-Arnold Convolutional Networks for Vision Recognition
by: Liu, Zhaoxiang, et al.
Published: (2026)
by: Liu, Zhaoxiang, et al.
Published: (2026)
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment
by: Li, Yantao, et al.
Published: (2026)
by: Li, Yantao, et al.
Published: (2026)
Unsupervised Industrial Anomaly Detection via Pattern Generative and Contrastive Networks
by: Huang, Jianfeng, et al.
Published: (2022)
by: Huang, Jianfeng, et al.
Published: (2022)
PhenoYieldNet: Learning Crop-Aware Phenological Responses for Multi-Crop Yield Prediction
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model
by: Wang, Yuqi, et al.
Published: (2024)
by: Wang, Yuqi, et al.
Published: (2024)
GateFuseNet: An Adaptive 3D Multimodal Neuroimaging Fusion Network for Parkinson's Disease Diagnosis
by: Jin, Rui, et al.
Published: (2025)
by: Jin, Rui, et al.
Published: (2025)
A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
by: Jiang, Siyang, et al.
Published: (2025)
by: Jiang, Siyang, et al.
Published: (2025)
RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation
by: Feng, Ganlin, et al.
Published: (2026)
by: Feng, Ganlin, et al.
Published: (2026)
PhysDrive: A Multimodal Remote Physiological Measurement Dataset for In-vehicle Driver Monitoring
by: Wang, Jiyao, et al.
Published: (2025)
by: Wang, Jiyao, et al.
Published: (2025)
EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs
by: Hu, He, et al.
Published: (2026)
by: Hu, He, et al.
Published: (2026)
PathInsight: Instruction Tuning of Multimodal Datasets and Models for Intelligence Assisted Diagnosis in Histopathology
by: Wu, Xiaomin, et al.
Published: (2024)
by: Wu, Xiaomin, et al.
Published: (2024)
MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction
by: Tang, Shuo, et al.
Published: (2025)
by: Tang, Shuo, et al.
Published: (2025)
Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation
by: Wang, Xinkun, et al.
Published: (2025)
by: Wang, Xinkun, et al.
Published: (2025)
Towards Practical Alzheimer's Disease Diagnosis: A Lightweight and Interpretable Spiking Neural Model
by: Wu, Changwei, et al.
Published: (2025)
by: Wu, Changwei, et al.
Published: (2025)
LoopNav: Benchmarking Spatial Consistency in World Models
by: Lian, Kewei, et al.
Published: (2025)
by: Lian, Kewei, et al.
Published: (2025)
Intelligent Pathological Diagnosis of Gestational Trophoblastic Diseases via Visual-Language Deep Learning Model
by: Liu, Yuhang, et al.
Published: (2026)
by: Liu, Yuhang, et al.
Published: (2026)
Emotion-LLaMAv2 and MMEVerse: A New Framework and Benchmark for Multimodal Emotion Understanding
by: Peng, Xiaojiang, et al.
Published: (2026)
by: Peng, Xiaojiang, et al.
Published: (2026)
EcoCropsAID: Economic Crops Aerial Image Dataset for Land Use Classification
by: Noppitak, Sangdaow, et al.
Published: (2024)
by: Noppitak, Sangdaow, et al.
Published: (2024)
MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
by: Zhang, Fan, et al.
Published: (2025)
by: Zhang, Fan, et al.
Published: (2025)
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
by: Liu, Zhihong, et al.
Published: (2026)
by: Liu, Zhihong, et al.
Published: (2026)
PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis
by: Hua, Shengyi, et al.
Published: (2025)
by: Hua, Shengyi, et al.
Published: (2025)
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis
by: Rui, Shaohao, et al.
Published: (2025)
by: Rui, Shaohao, et al.
Published: (2025)
MLLM-CL: Continual Learning for Multimodal Large Language Models
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
Similar Items
-
A Large Vision-Language Model based Environment Perception System for Visually Impaired People
by: Chen, Zezhou, et al.
Published: (2025) -
Hierarchical Deep Fusion Framework for Multi-dimensional Facial Forgery Detection -- The 2024 Global Deepfake Image Detection Challenge
by: Wang, Kohou, et al.
Published: (2025) -
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
by: Wang, Kohou, et al.
Published: (2025) -
MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
by: Zhao, Kaikai, et al.
Published: (2025) -
Patch-wise Auto-Encoder for Visual Anomaly Detection
by: Cui, Yajie, et al.
Published: (2023)