ViViD-5K: Vineyard vision dataset for field-based berry detection and segmentation and grape cluster closure estimation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916041474441216 |
|---|---|
| author | Tong, Xiangzhi Zhang, Chengrui Flaherty, Mac Garcia, Andre Matteo Gorman, Dominic Jaramillo, Jonathan Heuvel, Justine E. Vanden Jiang, Yu |
| author_facet | Tong, Xiangzhi Zhang, Chengrui Flaherty, Mac Garcia, Andre Matteo Gorman, Dominic Jaramillo, Jonathan Heuvel, Justine E. Vanden Jiang, Yu |
| contents | Cluster closure, defined as the progressive filling of gaps between the berries in a grape bunch, is a key trait in vineyard management, impacting disease risk. However, traditional visual scoring methods are labor-intensive, subjective, and lack temporal resolution. Existing datasets rarely support fine-grained berry-level analysis, limiting the development of robust deep learning models. In this work, we present ViViD-5k, a large-scale in-field Vineyard Vision Dataset containing 5,000 images with dense annotations, including over 648,000 berry centroids and cluster segmentation masks spanning 13 grape varieties. Building on this dataset, we introduce GrapeSAM, a two-stage visual pipeline that combines point-based berry localization with prompt-based segmentation using Segment Anything, followed by transformer-based cluster segmentation. The pipeline enables automated, in-field estimation of cluster closure with minimal supervision. Quantitative results demonstrate strong segmentation and counting accuracy across diverse conditions, while visualizations confirm robustness on both in-domain and out-of-domain samples. This work provides a scalable and objective alternative to manual compactness scoring and supports high-throughput grape phenotyping with enhanced spatial detail. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_24353 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ViViD-5K: Vineyard vision dataset for field-based berry detection and segmentation and grape cluster closure estimation Tong, Xiangzhi Zhang, Chengrui Flaherty, Mac Garcia, Andre Matteo Gorman, Dominic Jaramillo, Jonathan Heuvel, Justine E. Vanden Jiang, Yu Computer Vision and Pattern Recognition Other Quantitative Biology Cluster closure, defined as the progressive filling of gaps between the berries in a grape bunch, is a key trait in vineyard management, impacting disease risk. However, traditional visual scoring methods are labor-intensive, subjective, and lack temporal resolution. Existing datasets rarely support fine-grained berry-level analysis, limiting the development of robust deep learning models. In this work, we present ViViD-5k, a large-scale in-field Vineyard Vision Dataset containing 5,000 images with dense annotations, including over 648,000 berry centroids and cluster segmentation masks spanning 13 grape varieties. Building on this dataset, we introduce GrapeSAM, a two-stage visual pipeline that combines point-based berry localization with prompt-based segmentation using Segment Anything, followed by transformer-based cluster segmentation. The pipeline enables automated, in-field estimation of cluster closure with minimal supervision. Quantitative results demonstrate strong segmentation and counting accuracy across diverse conditions, while visualizations confirm robustness on both in-domain and out-of-domain samples. This work provides a scalable and objective alternative to manual compactness scoring and supports high-throughput grape phenotyping with enhanced spatial detail. |
| title | ViViD-5K: Vineyard vision dataset for field-based berry detection and segmentation and grape cluster closure estimation |
| topic | Computer Vision and Pattern Recognition Other Quantitative Biology |
| url | https://arxiv.org/abs/2605.24353 |