Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Weifeng, Wei, Xinyu, An, Ruichuan, Ren, Tianhe, Chen, Tingwei, Zhang, Renrui, Guo, Ziyu, Zhang, Wentao, Zhang, Lei, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Segment and Caption Anything
by: Huang, Xiaoke, et al.
Published: (2023)
by: Huang, Xiaoke, et al.
Published: (2023)
NTO3D: Neural Target Object 3D Reconstruction with Segment Anything
by: Wei, Xiaobao, et al.
Published: (2023)
by: Wei, Xiaobao, et al.
Published: (2023)
Matching Anything by Segmenting Anything
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
Rethinking VLM Representation for VLA Initialization
by: Lin, Weifeng, et al.
Published: (2026)
by: Lin, Weifeng, et al.
Published: (2026)
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
Detect Anything via Next Point Prediction
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
MedLSAM: Localize and Segment Anything Model for 3D CT Images
by: Lei, Wenhui, et al.
Published: (2023)
by: Lei, Wenhui, et al.
Published: (2023)
Segment Anything for Videos: A Systematic Survey
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
Biomedical SAM 2: Segment Anything in Biomedical Images and Videos
by: Yan, Zhiling, et al.
Published: (2024)
by: Yan, Zhiling, et al.
Published: (2024)
Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning
by: Shen, Zijun, et al.
Published: (2026)
by: Shen, Zijun, et al.
Published: (2026)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
by: Tang, Yunlong, et al.
Published: (2025)
by: Tang, Yunlong, et al.
Published: (2025)
Detect Anything 3D in the Wild
by: Zhang, Hanxue, et al.
Published: (2025)
by: Zhang, Hanxue, et al.
Published: (2025)
I-MedSAM: Implicit Medical Image Segmentation with Segment Anything
by: Wei, Xiaobao, et al.
Published: (2023)
by: Wei, Xiaobao, et al.
Published: (2023)
Segment Anything Is Not Always Perfect: An Investigation of SAM on Different Real-world Applications
by: Ji, Wei, et al.
Published: (2023)
by: Ji, Wei, et al.
Published: (2023)
Segment Anything in Medical Images
by: Ma, Jun, et al.
Published: (2023)
by: Ma, Jun, et al.
Published: (2023)
MESA: Matching Everything by Segmenting Anything
by: Zhang, Yesheng, et al.
Published: (2024)
by: Zhang, Yesheng, et al.
Published: (2024)
Matte Anything: Interactive Natural Image Matting with Segment Anything Models
by: Yao, Jingfeng, et al.
Published: (2023)
by: Yao, Jingfeng, et al.
Published: (2023)
UVOSAM: A Mask-free Paradigm for Unsupervised Video Object Segmentation via Segment Anything Model
by: Zhang, Zhenghao, et al.
Published: (2023)
by: Zhang, Zhenghao, et al.
Published: (2023)
Segment Anything, Even Occluded
by: Tai, Wei-En, et al.
Published: (2025)
by: Tai, Wei-En, et al.
Published: (2025)
SAM 2: Segment Anything in Images and Videos
by: Ravi, Nikhila, et al.
Published: (2024)
by: Ravi, Nikhila, et al.
Published: (2024)
Deshadow-Anything: When Segment Anything Model Meets Zero-shot shadow removal
by: Zhang, Xiao Feng, et al.
Published: (2023)
by: Zhang, Xiao Feng, et al.
Published: (2023)
Enhancing the Reliability of Segment Anything Model for Auto-Prompting Medical Image Segmentation with Uncertainty Rectification
by: Zhang, Yichi, et al.
Published: (2023)
by: Zhang, Yichi, et al.
Published: (2023)
SAM3-UNet: Simplified Adaptation of Segment Anything Model 3
by: Xiong, Xinyu, et al.
Published: (2025)
by: Xiong, Xinyu, et al.
Published: (2025)
Tracking and Segmenting Anything in Any Modality
by: Zhang, Tianlu, et al.
Published: (2025)
by: Zhang, Tianlu, et al.
Published: (2025)
Segment Anything in Medical Images and Videos: Benchmark and Deployment
by: Ma, Jun, et al.
Published: (2024)
by: Ma, Jun, et al.
Published: (2024)
AffordanceSAM: Segment Anything Once More in Affordance Grounding
by: Jiang, Dengyang, et al.
Published: (2025)
by: Jiang, Dengyang, et al.
Published: (2025)
Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
by: Xu, Guoping, et al.
Published: (2025)
by: Xu, Guoping, et al.
Published: (2025)
Segment Anything Model for Brain Tumor Segmentation
by: Zhang, Peng, et al.
Published: (2023)
by: Zhang, Peng, et al.
Published: (2023)
URECA: Unique Region Caption Anything
by: Lim, Sangbeom, et al.
Published: (2025)
by: Lim, Sangbeom, et al.
Published: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
SAM3-I: Segment Anything with Instructions
by: Li, Jingjing, et al.
Published: (2025)
by: Li, Jingjing, et al.
Published: (2025)
Segment Anything in 3D with Radiance Fields
by: Cen, Jiazhong, et al.
Published: (2023)
by: Cen, Jiazhong, et al.
Published: (2023)
Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis
by: Yuan, Cheng, et al.
Published: (2024)
by: Yuan, Cheng, et al.
Published: (2024)
Learning to Prompt Segment Anything Models
by: Huang, Jiaxing, et al.
Published: (2024)
by: Huang, Jiaxing, et al.
Published: (2024)
AnimateAnything: Consistent and Controllable Animation for Video Generation
by: Lei, Guojun, et al.
Published: (2024)
by: Lei, Guojun, et al.
Published: (2024)
Register Anything: Estimating "Corresponding Prompts" for Segment Anything Model
by: Huang, Shiqi, et al.
Published: (2025)
by: Huang, Shiqi, et al.
Published: (2025)
IRSAM: Advancing Segment Anything Model for Infrared Small Target Detection
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
GENIUS: Generative Fluid Intelligence Evaluation Suite
by: An, Ruichuan, et al.
Published: (2026)
by: An, Ruichuan, et al.
Published: (2026)
Describe Anything in Medical Images
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
Segment Anything in Pathology Images with Natural Language
by: Chen, Zhixuan, et al.
Published: (2025)
by: Chen, Zhixuan, et al.
Published: (2025)
Similar Items
-
Segment and Caption Anything
by: Huang, Xiaoke, et al.
Published: (2023) -
NTO3D: Neural Target Object 3D Reconstruction with Segment Anything
by: Wei, Xiaobao, et al.
Published: (2023) -
Matching Anything by Segmenting Anything
by: Li, Siyuan, et al.
Published: (2024) -
Rethinking VLM Representation for VLA Initialization
by: Lin, Weifeng, et al.
Published: (2026) -
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)