Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xing, Junhao, Miyakawa, Ryohei, Yang, Yang, Liu, Xinpeng, Shinoda, Risa, Santo, Hiroaki, Toda, Yosuke, Okura, Fumio
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911327232983040
author Xing, Junhao
Miyakawa, Ryohei
Yang, Yang
Liu, Xinpeng
Shinoda, Risa
Santo, Hiroaki
Toda, Yosuke
Okura, Fumio
author_facet Xing, Junhao
Miyakawa, Ryohei
Yang, Yang
Liu, Xinpeng
Shinoda, Risa
Santo, Hiroaki
Toda, Yosuke
Okura, Fumio
contents Foundation segmentation models achieve reasonable leaf instance extraction from top-view crop images without training (i.e., zero-shot). However, segmenting entire plant individuals with each consisting of multiple overlapping leaves remains challenging. This problem is referred to as a hierarchical segmentation task, typically requiring annotated training datasets, which are often species-specific and require notable human labor. To address this, we introduce ZeroPlantSeg, a zero-shot segmentation for rosette-shaped plant individuals from top-view images. We integrate a foundation segmentation model, extracting leaf instances, and a vision-language model, reasoning about plants' structures to extract plant individuals without additional training. Evaluations on datasets with multiple plant species, growth stages, and shooting environments demonstrate that our method surpasses existing zero-shot methods and achieves better cross-domain performance than supervised methods. Implementations are available at https://github.com/JunhaoXing/ZeroPlantSeg.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09116
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention
Xing, Junhao
Miyakawa, Ryohei
Yang, Yang
Liu, Xinpeng
Shinoda, Risa
Santo, Hiroaki
Toda, Yosuke
Okura, Fumio
Computer Vision and Pattern Recognition
Foundation segmentation models achieve reasonable leaf instance extraction from top-view crop images without training (i.e., zero-shot). However, segmenting entire plant individuals with each consisting of multiple overlapping leaves remains challenging. This problem is referred to as a hierarchical segmentation task, typically requiring annotated training datasets, which are often species-specific and require notable human labor. To address this, we introduce ZeroPlantSeg, a zero-shot segmentation for rosette-shaped plant individuals from top-view images. We integrate a foundation segmentation model, extracting leaf instances, and a vision-language model, reasoning about plants' structures to extract plant individuals without additional training. Evaluations on datasets with multiple plant species, growth stages, and shooting environments demonstrate that our method surpasses existing zero-shot methods and achieves better cross-domain performance than supervised methods. Implementations are available at https://github.com/JunhaoXing/ZeroPlantSeg.
title Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.09116