Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915328549715968 |
|---|---|
| author | Lin, Weifeng Wei, Xinyu An, Ruichuan Ren, Tianhe Chen, Tingwei Zhang, Renrui Guo, Ziyu Zhang, Wentao Zhang, Lei Li, Hongsheng |
| author_facet | Lin, Weifeng Wei, Xinyu An, Ruichuan Ren, Tianhe Chen, Tingwei Zhang, Renrui Guo, Ziyu Zhang, Wentao Zhang, Lei Li, Hongsheng |
| contents | We present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our approach extends the powerful segmentation model SAM 2 by integrating Large Language Models (LLMs), enabling simultaneous object segmentation with the generation of diverse, region-specific semantic outputs, including categories, label definition, functional explanations, and detailed captions. A key component, Semantic Perceiver, is introduced to efficiently transform SAM 2's rich visual features, which inherently carry general vision, localization, and semantic priors into multi-modal tokens for LLM comprehension. To support robust multi-granularity understanding, we also develop a dedicated data refinement and augmentation pipeline, yielding a high-quality dataset of 1.5M image and 0.6M video region-semantic annotations, including novel region-level streaming video caption data. PAM is designed for lightweightness and efficiency, while also demonstrates strong performance across a diverse range of region understanding tasks. It runs 1.2-2.4x faster and consumes less GPU memory than prior approaches, offering a practical solution for real-world applications. We believe that our effective approach will serve as a strong baseline for future research in region-level visual understanding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_05302 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Lin, Weifeng Wei, Xinyu An, Ruichuan Ren, Tianhe Chen, Tingwei Zhang, Renrui Guo, Ziyu Zhang, Wentao Zhang, Lei Li, Hongsheng Computer Vision and Pattern Recognition We present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our approach extends the powerful segmentation model SAM 2 by integrating Large Language Models (LLMs), enabling simultaneous object segmentation with the generation of diverse, region-specific semantic outputs, including categories, label definition, functional explanations, and detailed captions. A key component, Semantic Perceiver, is introduced to efficiently transform SAM 2's rich visual features, which inherently carry general vision, localization, and semantic priors into multi-modal tokens for LLM comprehension. To support robust multi-granularity understanding, we also develop a dedicated data refinement and augmentation pipeline, yielding a high-quality dataset of 1.5M image and 0.6M video region-semantic annotations, including novel region-level streaming video caption data. PAM is designed for lightweightness and efficiency, while also demonstrates strong performance across a diverse range of region understanding tasks. It runs 1.2-2.4x faster and consumes less GPU memory than prior approaches, offering a practical solution for real-world applications. We believe that our effective approach will serve as a strong baseline for future research in region-level visual understanding. |
| title | Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.05302 |