Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Weifeng, Wei, Xinyu, An, Ruichuan, Ren, Tianhe, Chen, Tingwei, Zhang, Renrui, Guo, Ziyu, Zhang, Wentao, Zhang, Lei, Li, Hongsheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915328549715968
author Lin, Weifeng
Wei, Xinyu
An, Ruichuan
Ren, Tianhe
Chen, Tingwei
Zhang, Renrui
Guo, Ziyu
Zhang, Wentao
Zhang, Lei
Li, Hongsheng
author_facet Lin, Weifeng
Wei, Xinyu
An, Ruichuan
Ren, Tianhe
Chen, Tingwei
Zhang, Renrui
Guo, Ziyu
Zhang, Wentao
Zhang, Lei
Li, Hongsheng
contents We present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our approach extends the powerful segmentation model SAM 2 by integrating Large Language Models (LLMs), enabling simultaneous object segmentation with the generation of diverse, region-specific semantic outputs, including categories, label definition, functional explanations, and detailed captions. A key component, Semantic Perceiver, is introduced to efficiently transform SAM 2's rich visual features, which inherently carry general vision, localization, and semantic priors into multi-modal tokens for LLM comprehension. To support robust multi-granularity understanding, we also develop a dedicated data refinement and augmentation pipeline, yielding a high-quality dataset of 1.5M image and 0.6M video region-semantic annotations, including novel region-level streaming video caption data. PAM is designed for lightweightness and efficiency, while also demonstrates strong performance across a diverse range of region understanding tasks. It runs 1.2-2.4x faster and consumes less GPU memory than prior approaches, offering a practical solution for real-world applications. We believe that our effective approach will serve as a strong baseline for future research in region-level visual understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05302
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Lin, Weifeng
Wei, Xinyu
An, Ruichuan
Ren, Tianhe
Chen, Tingwei
Zhang, Renrui
Guo, Ziyu
Zhang, Wentao
Zhang, Lei
Li, Hongsheng
Computer Vision and Pattern Recognition
We present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our approach extends the powerful segmentation model SAM 2 by integrating Large Language Models (LLMs), enabling simultaneous object segmentation with the generation of diverse, region-specific semantic outputs, including categories, label definition, functional explanations, and detailed captions. A key component, Semantic Perceiver, is introduced to efficiently transform SAM 2's rich visual features, which inherently carry general vision, localization, and semantic priors into multi-modal tokens for LLM comprehension. To support robust multi-granularity understanding, we also develop a dedicated data refinement and augmentation pipeline, yielding a high-quality dataset of 1.5M image and 0.6M video region-semantic annotations, including novel region-level streaming video caption data. PAM is designed for lightweightness and efficiency, while also demonstrates strong performance across a diverse range of region understanding tasks. It runs 1.2-2.4x faster and consumes less GPU memory than prior approaches, offering a practical solution for real-world applications. We believe that our effective approach will serve as a strong baseline for future research in region-level visual understanding.
title Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.05302