CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Wentao, Fang, Tao, Lu, Lina, Wang, Lifei, Zhong, Weihe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Agri-CPJ: A Training-Free Explainable Framework for Agricultural Pest Diagnosis Using Caption-Prompt-Judge and LLM-as-a-Judge
por: Zhang, Wentao, et al.
Publicado: (2026)
por: Zhang, Wentao, et al.
Publicado: (2026)
Agri-R1: Agricultural Reasoning for Disease Diagnosis via Automated-Synthesis and Reinforcement Learning
por: Zhang, Wentao, et al.
Publicado: (2026)
por: Zhang, Wentao, et al.
Publicado: (2026)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
por: Matsuda, Kazuki, et al.
Publicado: (2025)
por: Matsuda, Kazuki, et al.
Publicado: (2025)
Judge Anything: MLLM as a Judge Across Any Modality
por: Pu, Shu, et al.
Publicado: (2025)
por: Pu, Shu, et al.
Publicado: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
por: Chen, Dongping, et al.
Publicado: (2024)
por: Chen, Dongping, et al.
Publicado: (2024)
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
por: Waheed, Abdul, et al.
Publicado: (2025)
por: Waheed, Abdul, et al.
Publicado: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
por: Laskar, Md Tahmid Rahman, et al.
Publicado: (2025)
por: Laskar, Md Tahmid Rahman, et al.
Publicado: (2025)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
por: Lee, Sua, et al.
Publicado: (2026)
por: Lee, Sua, et al.
Publicado: (2026)
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
por: Xie, Yupeng, et al.
Publicado: (2025)
por: Xie, Yupeng, et al.
Publicado: (2025)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
por: Hwang, Yerin, et al.
Publicado: (2025)
por: Hwang, Yerin, et al.
Publicado: (2025)
A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
por: Betala, Siddharth, et al.
Publicado: (2025)
por: Betala, Siddharth, et al.
Publicado: (2025)
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
por: Ding, Meidan, et al.
Publicado: (2025)
por: Ding, Meidan, et al.
Publicado: (2025)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
por: Zhao, Yi, et al.
Publicado: (2026)
por: Zhao, Yi, et al.
Publicado: (2026)
OmniCaptioner: One Captioner to Rule Them All
por: Lu, Yiting, et al.
Publicado: (2025)
por: Lu, Yiting, et al.
Publicado: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
por: Liu, Zheng, et al.
Publicado: (2024)
por: Liu, Zheng, et al.
Publicado: (2024)
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
por: Liu, Runzhou, et al.
Publicado: (2026)
por: Liu, Runzhou, et al.
Publicado: (2026)
Multi-LLM Collaborative Caption Generation in Scientific Documents
por: Kim, Jaeyoung, et al.
Publicado: (2025)
por: Kim, Jaeyoung, et al.
Publicado: (2025)
Computer-Use Agents as Judges for Generative User Interface
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
PromptMRG: Diagnosis-Driven Prompts for Medical Report Generation
por: Jin, Haibo, et al.
Publicado: (2023)
por: Jin, Haibo, et al.
Publicado: (2023)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
por: Cheng, Kanzhi, et al.
Publicado: (2025)
por: Cheng, Kanzhi, et al.
Publicado: (2025)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
por: Kim, Hyunjong, et al.
Publicado: (2025)
por: Kim, Hyunjong, et al.
Publicado: (2025)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
por: Ma, Ziyu, et al.
Publicado: (2024)
por: Ma, Ziyu, et al.
Publicado: (2024)
Image Captioning via Compact Bidirectional Architecture
por: Song, Zijie, et al.
Publicado: (2022)
por: Song, Zijie, et al.
Publicado: (2022)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
por: Xu, Run, et al.
Publicado: (2026)
por: Xu, Run, et al.
Publicado: (2026)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
por: Freitas, Miguel Monte e, et al.
Publicado: (2026)
por: Freitas, Miguel Monte e, et al.
Publicado: (2026)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
por: Li, Yuying, et al.
Publicado: (2025)
por: Li, Yuying, et al.
Publicado: (2025)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
por: Ji, Yatai, et al.
Publicado: (2024)
por: Ji, Yatai, et al.
Publicado: (2024)
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
por: Zhou, Qiji, et al.
Publicado: (2024)
por: Zhou, Qiji, et al.
Publicado: (2024)
MLLM-as-a-Judge for Image Safety without Human Labeling
por: Wang, Zhenting, et al.
Publicado: (2024)
por: Wang, Zhenting, et al.
Publicado: (2024)
Altogether: Image Captioning via Re-aligning Alt-text
por: Xu, Hu, et al.
Publicado: (2024)
por: Xu, Hu, et al.
Publicado: (2024)
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
por: Diesendruck, Maurice, et al.
Publicado: (2024)
por: Diesendruck, Maurice, et al.
Publicado: (2024)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
por: Xing, Long, et al.
Publicado: (2025)
por: Xing, Long, et al.
Publicado: (2025)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
por: Kumar, Divake, et al.
Publicado: (2026)
por: Kumar, Divake, et al.
Publicado: (2026)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
por: Chaffin, Antoine, et al.
Publicado: (2024)
por: Chaffin, Antoine, et al.
Publicado: (2024)
A CNN-Based Malaria Diagnosis from Blood Cell Images with SHAP and LIME Explainability
por: Abir, Md. Ismiel Hossen, et al.
Publicado: (2025)
por: Abir, Md. Ismiel Hossen, et al.
Publicado: (2025)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
por: Chen, Zhaorun, et al.
Publicado: (2024)
por: Chen, Zhaorun, et al.
Publicado: (2024)
PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
por: Ananthram, Amith, et al.
Publicado: (2025)
por: Ananthram, Amith, et al.
Publicado: (2025)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
por: Wu, Xuyang, et al.
Publicado: (2024)
por: Wu, Xuyang, et al.
Publicado: (2024)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
por: Ozaki, Shintaro, et al.
Publicado: (2025)
por: Ozaki, Shintaro, et al.
Publicado: (2025)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
por: Singh, Ayush, et al.
Publicado: (2024)
por: Singh, Ayush, et al.
Publicado: (2024)
Ejemplares similares
-
Agri-CPJ: A Training-Free Explainable Framework for Agricultural Pest Diagnosis Using Caption-Prompt-Judge and LLM-as-a-Judge
por: Zhang, Wentao, et al.
Publicado: (2026) -
Agri-R1: Agricultural Reasoning for Disease Diagnosis via Automated-Synthesis and Reinforcement Learning
por: Zhang, Wentao, et al.
Publicado: (2026) -
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
por: Matsuda, Kazuki, et al.
Publicado: (2025) -
Judge Anything: MLLM as a Judge Across Any Modality
por: Pu, Shu, et al.
Publicado: (2025) -
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
por: Chen, Dongping, et al.
Publicado: (2024)