PicoSAM2: Low-Latency Segmentation In-Sensor for Edge Vision Applications
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908642753642496 |
|---|---|
| author | Bonazzi, Pietro Farronato, Nicola Zihlmann, Stefan Qin, Haotong Magno, Michele |
| author_facet | Bonazzi, Pietro Farronato, Nicola Zihlmann, Stefan Qin, Haotong Magno, Michele |
| contents | Real-time, on-device segmentation is critical for latency-sensitive and privacy-aware applications like smart glasses and IoT devices. We introduce PicoSAM2, a lightweight (1.3M parameters, 336M MACs) promptable segmentation model optimized for edge and in-sensor execution, including the Sony IMX500. It builds on a depthwise separable U-Net, with knowledge distillation and fixed-point prompt encoding to learn from the Segment Anything Model 2 (SAM2). On COCO and LVIS, it achieves 51.9% and 44.9% mIoU, respectively. The quantized model (1.22MB) runs at 14.3 ms on the IMX500-achieving 86 MACs/cycle, making it the only model meeting both memory and compute constraints for in-sensor deployment. Distillation boosts LVIS performance by +3.5% mIoU and +5.1% mAP. These results demonstrate that efficient, promptable segmentation is feasible directly on-camera, enabling privacy-preserving vision without cloud or host processing. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_18807 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | PicoSAM2: Low-Latency Segmentation In-Sensor for Edge Vision Applications Bonazzi, Pietro Farronato, Nicola Zihlmann, Stefan Qin, Haotong Magno, Michele Computer Vision and Pattern Recognition Real-time, on-device segmentation is critical for latency-sensitive and privacy-aware applications like smart glasses and IoT devices. We introduce PicoSAM2, a lightweight (1.3M parameters, 336M MACs) promptable segmentation model optimized for edge and in-sensor execution, including the Sony IMX500. It builds on a depthwise separable U-Net, with knowledge distillation and fixed-point prompt encoding to learn from the Segment Anything Model 2 (SAM2). On COCO and LVIS, it achieves 51.9% and 44.9% mIoU, respectively. The quantized model (1.22MB) runs at 14.3 ms on the IMX500-achieving 86 MACs/cycle, making it the only model meeting both memory and compute constraints for in-sensor deployment. Distillation boosts LVIS performance by +3.5% mIoU and +5.1% mAP. These results demonstrate that efficient, promptable segmentation is feasible directly on-camera, enabling privacy-preserving vision without cloud or host processing. |
| title | PicoSAM2: Low-Latency Segmentation In-Sensor for Edge Vision Applications |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.18807 |