ViViD-5K: Vineyard vision dataset for field-based berry detection and segmentation and grape cluster closure estimation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tong, Xiangzhi, Zhang, Chengrui, Flaherty, Mac, Garcia, Andre Matteo, Gorman, Dominic, Jaramillo, Jonathan, Heuvel, Justine E. Vanden, Jiang, Yu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916041474441216
author Tong, Xiangzhi
Zhang, Chengrui
Flaherty, Mac
Garcia, Andre Matteo
Gorman, Dominic
Jaramillo, Jonathan
Heuvel, Justine E. Vanden
Jiang, Yu
author_facet Tong, Xiangzhi
Zhang, Chengrui
Flaherty, Mac
Garcia, Andre Matteo
Gorman, Dominic
Jaramillo, Jonathan
Heuvel, Justine E. Vanden
Jiang, Yu
contents Cluster closure, defined as the progressive filling of gaps between the berries in a grape bunch, is a key trait in vineyard management, impacting disease risk. However, traditional visual scoring methods are labor-intensive, subjective, and lack temporal resolution. Existing datasets rarely support fine-grained berry-level analysis, limiting the development of robust deep learning models. In this work, we present ViViD-5k, a large-scale in-field Vineyard Vision Dataset containing 5,000 images with dense annotations, including over 648,000 berry centroids and cluster segmentation masks spanning 13 grape varieties. Building on this dataset, we introduce GrapeSAM, a two-stage visual pipeline that combines point-based berry localization with prompt-based segmentation using Segment Anything, followed by transformer-based cluster segmentation. The pipeline enables automated, in-field estimation of cluster closure with minimal supervision. Quantitative results demonstrate strong segmentation and counting accuracy across diverse conditions, while visualizations confirm robustness on both in-domain and out-of-domain samples. This work provides a scalable and objective alternative to manual compactness scoring and supports high-throughput grape phenotyping with enhanced spatial detail.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24353
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ViViD-5K: Vineyard vision dataset for field-based berry detection and segmentation and grape cluster closure estimation
Tong, Xiangzhi
Zhang, Chengrui
Flaherty, Mac
Garcia, Andre Matteo
Gorman, Dominic
Jaramillo, Jonathan
Heuvel, Justine E. Vanden
Jiang, Yu
Computer Vision and Pattern Recognition
Other Quantitative Biology
Cluster closure, defined as the progressive filling of gaps between the berries in a grape bunch, is a key trait in vineyard management, impacting disease risk. However, traditional visual scoring methods are labor-intensive, subjective, and lack temporal resolution. Existing datasets rarely support fine-grained berry-level analysis, limiting the development of robust deep learning models. In this work, we present ViViD-5k, a large-scale in-field Vineyard Vision Dataset containing 5,000 images with dense annotations, including over 648,000 berry centroids and cluster segmentation masks spanning 13 grape varieties. Building on this dataset, we introduce GrapeSAM, a two-stage visual pipeline that combines point-based berry localization with prompt-based segmentation using Segment Anything, followed by transformer-based cluster segmentation. The pipeline enables automated, in-field estimation of cluster closure with minimal supervision. Quantitative results demonstrate strong segmentation and counting accuracy across diverse conditions, while visualizations confirm robustness on both in-domain and out-of-domain samples. This work provides a scalable and objective alternative to manual compactness scoring and supports high-throughput grape phenotyping with enhanced spatial detail.
title ViViD-5K: Vineyard vision dataset for field-based berry detection and segmentation and grape cluster closure estimation
topic Computer Vision and Pattern Recognition
Other Quantitative Biology
url https://arxiv.org/abs/2605.24353