Vision Meets Language: A RAG-Augmented YOLOv8 Framework for Coffee Disease Diagnosis and Farmer Assistance
Fuente:
arXiv
Salvato in:
| Autore principale: | Mondal, Semanto |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Malayalam Sign Language Identification using Finetuned YOLOv8 and Computer Vision Techniques
di: K., Abhinand, et al.
Pubblicazione: (2024)
di: K., Abhinand, et al.
Pubblicazione: (2024)
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
di: Pellegrini, Chantal, et al.
Pubblicazione: (2023)
di: Pellegrini, Chantal, et al.
Pubblicazione: (2023)
SmoGVLM: A Small, Graph-enhanced Vision-Language Model
di: Mondal, Debjyoti, et al.
Pubblicazione: (2026)
di: Mondal, Debjyoti, et al.
Pubblicazione: (2026)
Improved YOLOv12 with LLM-Generated Synthetic Data for Enhanced Apple Detection and Benchmarking Against YOLOv11 and YOLOv10
di: Sapkota, Ranjan, et al.
Pubblicazione: (2025)
di: Sapkota, Ranjan, et al.
Pubblicazione: (2025)
A Two-Stage Multitask Vision-Language Framework for Explainable Crop Disease Visual Question Answering
di: Hossain, Md. Zahid, et al.
Pubblicazione: (2026)
di: Hossain, Md. Zahid, et al.
Pubblicazione: (2026)
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
di: Xia, Peng, et al.
Pubblicazione: (2024)
di: Xia, Peng, et al.
Pubblicazione: (2024)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
di: Lokesh, K, et al.
Pubblicazione: (2026)
di: Lokesh, K, et al.
Pubblicazione: (2026)
YOLOv5, YOLOv8 and YOLOv10: The Go-To Detectors for Real-time Vision
di: Hussain, Muhammad
Pubblicazione: (2024)
di: Hussain, Muhammad
Pubblicazione: (2024)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
di: Loo, Gowen, et al.
Pubblicazione: (2025)
di: Loo, Gowen, et al.
Pubblicazione: (2025)
Generalizable Entity Grounding via Assistance of Large Language Model
di: Qi, Lu, et al.
Pubblicazione: (2024)
di: Qi, Lu, et al.
Pubblicazione: (2024)
Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit
di: Chen, Qizhou, et al.
Pubblicazione: (2024)
di: Chen, Qizhou, et al.
Pubblicazione: (2024)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
di: Qian, Kangan, et al.
Pubblicazione: (2025)
di: Qian, Kangan, et al.
Pubblicazione: (2025)
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
LaB-RAG: Label Boosted Retrieval Augmented Generation for Radiology Report Generation
di: Song, Steven, et al.
Pubblicazione: (2024)
di: Song, Steven, et al.
Pubblicazione: (2024)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
di: Chang, Yue, et al.
Pubblicazione: (2024)
di: Chang, Yue, et al.
Pubblicazione: (2024)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
di: Wu, Yin, et al.
Pubblicazione: (2025)
di: Wu, Yin, et al.
Pubblicazione: (2025)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
di: Fan, Zhiwen, et al.
Pubblicazione: (2025)
di: Fan, Zhiwen, et al.
Pubblicazione: (2025)
Evaluating Data Augmentation Techniques for Coffee Leaf Disease Classification
di: Gheorghiu, Adrian, et al.
Pubblicazione: (2024)
di: Gheorghiu, Adrian, et al.
Pubblicazione: (2024)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
di: Sun, Yubo, et al.
Pubblicazione: (2025)
di: Sun, Yubo, et al.
Pubblicazione: (2025)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
di: Wang, Qiuchen, et al.
Pubblicazione: (2026)
di: Wang, Qiuchen, et al.
Pubblicazione: (2026)
Trajectory Prediction Meets Large Language Models: A Survey
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
di: Zhao, Yi, et al.
Pubblicazione: (2026)
di: Zhao, Yi, et al.
Pubblicazione: (2026)
Insight: A Multi-Modal Diagnostic Pipeline using LLMs for Ocular Surface Disease Diagnosis
di: Yeh, Chun-Hsiao, et al.
Pubblicazione: (2024)
di: Yeh, Chun-Hsiao, et al.
Pubblicazione: (2024)
A Comparative Analysis of YOLOv5, YOLOv8, and YOLOv10 in Kitchen Safety
di: Geetha, Athulya Sundaresan, et al.
Pubblicazione: (2024)
di: Geetha, Athulya Sundaresan, et al.
Pubblicazione: (2024)
Language Models Meet Anomaly Detection for Better Interpretability and Generalizability
di: Li, Jun, et al.
Pubblicazione: (2024)
di: Li, Jun, et al.
Pubblicazione: (2024)
Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
di: Li, Zhecheng, et al.
Pubblicazione: (2025)
di: Li, Zhecheng, et al.
Pubblicazione: (2025)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
di: Sathe, Ashutosh, et al.
Pubblicazione: (2024)
di: Sathe, Ashutosh, et al.
Pubblicazione: (2024)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
di: Ding, Yi, et al.
Pubblicazione: (2026)
di: Ding, Yi, et al.
Pubblicazione: (2026)
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
Agri-R1: Agricultural Reasoning for Disease Diagnosis via Automated-Synthesis and Reinforcement Learning
di: Zhang, Wentao, et al.
Pubblicazione: (2026)
di: Zhang, Wentao, et al.
Pubblicazione: (2026)
A Comparative Study of YOLOv8 to YOLOv11 Performance in Underwater Vision Tasks
di: Hung, Gordon, et al.
Pubblicazione: (2025)
di: Hung, Gordon, et al.
Pubblicazione: (2025)
Comprehensive Performance Evaluation of YOLOv11, YOLOv10, YOLOv9, YOLOv8 and YOLOv5 on Object Detection of Power Equipment
di: He, Zijian, et al.
Pubblicazione: (2024)
di: He, Zijian, et al.
Pubblicazione: (2024)
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
di: Hu, Chan-Wei, et al.
Pubblicazione: (2025)
di: Hu, Chan-Wei, et al.
Pubblicazione: (2025)
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
di: Zhao, Beidi, et al.
Pubblicazione: (2026)
di: Zhao, Beidi, et al.
Pubblicazione: (2026)
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
di: Cai, Hengxing, et al.
Pubblicazione: (2025)
di: Cai, Hengxing, et al.
Pubblicazione: (2025)
Lightweight Shrimp Disease Detection Research Based on YOLOv8n
di: Yuhuan, Fei, et al.
Pubblicazione: (2025)
di: Yuhuan, Fei, et al.
Pubblicazione: (2025)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
di: Ye, Jiacheng, et al.
Pubblicazione: (2025)
di: Ye, Jiacheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Malayalam Sign Language Identification using Finetuned YOLOv8 and Computer Vision Techniques
di: K., Abhinand, et al.
Pubblicazione: (2024) -
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
di: Wang, Wenbin, et al.
Pubblicazione: (2025) -
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
di: Pellegrini, Chantal, et al.
Pubblicazione: (2023) -
SmoGVLM: A Small, Graph-enhanced Vision-Language Model
di: Mondal, Debjyoti, et al.
Pubblicazione: (2026) -
Improved YOLOv12 with LLM-Generated Synthetic Data for Enhanced Apple Detection and Benchmarking Against YOLOv11 and YOLOv10
di: Sapkota, Ranjan, et al.
Pubblicazione: (2025)