P2Object: Single Point Supervised Object Detection and Instance Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Pengfei, Yu, Xuehui, Han, Xumeng, Wang, Kuiran, Li, Guorong, Xie, Lingxi, Han, Zhenjun, Jiao, Jianbin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909574044319744
author Chen, Pengfei
Yu, Xuehui
Han, Xumeng
Wang, Kuiran
Li, Guorong
Xie, Lingxi
Han, Zhenjun
Jiao, Jianbin
author_facet Chen, Pengfei
Yu, Xuehui
Han, Xumeng
Wang, Kuiran
Li, Guorong
Xie, Lingxi
Han, Zhenjun
Jiao, Jianbin
contents Object recognition using single-point supervision has attracted increasing attention recently. However, the performance gap compared with fully-supervised algorithms remains large. Previous works generated class-agnostic \textbf{\textit{proposals in an image}} offline and then treated mixed candidates as a single bag, putting a huge burden on multiple instance learning (MIL). In this paper, we introduce Point-to-Box Network (P2BNet), which constructs balanced \textbf{\textit{instance-level proposal bags}} by generating proposals in an anchor-like way and refining the proposals in a coarse-to-fine paradigm. Through further research, we find that the bag of proposals, either at the image level or the instance level, is established on discrete box sampling. This leads the pseudo box estimation into a sub-optimal solution, resulting in the truncation of object boundaries or the excessive inclusion of background. Hence, we conduct a series exploration of discrete-to-continuous optimization, yielding P2BNet++ and Point-to-Mask Network (P2MNet). P2BNet++ conducts an approximately continuous proposal sampling strategy by better utilizing spatial clues. P2MNet further introduces low-level image information to assist in pixel prediction, and a boundary self-prediction is designed to relieve the limitation of the estimated boxes. Benefiting from the continuous object-aware \textbf{\textit{pixel-level perception}}, P2MNet can generate more precise bounding boxes and generalize to segmentation tasks. Our method largely surpasses the previous methods in terms of the mean average precision on COCO, VOC, SBD, and Cityscapes, demonstrating great potential to bridge the performance gap compared with fully supervised tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2504_07813
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle P2Object: Single Point Supervised Object Detection and Instance Segmentation
Chen, Pengfei
Yu, Xuehui
Han, Xumeng
Wang, Kuiran
Li, Guorong
Xie, Lingxi
Han, Zhenjun
Jiao, Jianbin
Computer Vision and Pattern Recognition
Object recognition using single-point supervision has attracted increasing attention recently. However, the performance gap compared with fully-supervised algorithms remains large. Previous works generated class-agnostic \textbf{\textit{proposals in an image}} offline and then treated mixed candidates as a single bag, putting a huge burden on multiple instance learning (MIL). In this paper, we introduce Point-to-Box Network (P2BNet), which constructs balanced \textbf{\textit{instance-level proposal bags}} by generating proposals in an anchor-like way and refining the proposals in a coarse-to-fine paradigm. Through further research, we find that the bag of proposals, either at the image level or the instance level, is established on discrete box sampling. This leads the pseudo box estimation into a sub-optimal solution, resulting in the truncation of object boundaries or the excessive inclusion of background. Hence, we conduct a series exploration of discrete-to-continuous optimization, yielding P2BNet++ and Point-to-Mask Network (P2MNet). P2BNet++ conducts an approximately continuous proposal sampling strategy by better utilizing spatial clues. P2MNet further introduces low-level image information to assist in pixel prediction, and a boundary self-prediction is designed to relieve the limitation of the estimated boxes. Benefiting from the continuous object-aware \textbf{\textit{pixel-level perception}}, P2MNet can generate more precise bounding boxes and generalize to segmentation tasks. Our method largely surpasses the previous methods in terms of the mean average precision on COCO, VOC, SBD, and Cityscapes, demonstrating great potential to bridge the performance gap compared with fully supervised tasks.
title P2Object: Single Point Supervised Object Detection and Instance Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.07813