_version_ 1866910080846266368
author Carion, Nicolas
Gustafson, Laura
Hu, Yuan-Ting
Debnath, Shoubhik
Hu, Ronghang
Suris, Didac
Ryali, Chaitanya
Alwala, Kalyan Vasudev
Khedr, Haitham
Huang, Andrew
Lei, Jie
Ma, Tengyu
Guo, Baishan
Kalla, Arpit
Marks, Markus
Greer, Joseph
Wang, Meng
Sun, Peize
Rädle, Roman
Afouras, Triantafyllos
Mavroudi, Effrosyni
Xu, Katherine
Wu, Tsung-Han
Zhou, Yu
Momeni, Liliane
Hazra, Rishi
Ding, Shuangrui
Vaze, Sagar
Porcher, Francois
Li, Feng
Li, Siyuan
Kamath, Aishwarya
Cheng, Ho Kei
Dollár, Piotr
Ravi, Nikhila
Saenko, Kate
Zhang, Pengchuan
Feichtenhofer, Christoph
author_facet Carion, Nicolas
Gustafson, Laura
Hu, Yuan-Ting
Debnath, Shoubhik
Hu, Ronghang
Suris, Didac
Ryali, Chaitanya
Alwala, Kalyan Vasudev
Khedr, Haitham
Huang, Andrew
Lei, Jie
Ma, Tengyu
Guo, Baishan
Kalla, Arpit
Marks, Markus
Greer, Joseph
Wang, Meng
Sun, Peize
Rädle, Roman
Afouras, Triantafyllos
Mavroudi, Effrosyni
Xu, Katherine
Wu, Tsung-Han
Zhou, Yu
Momeni, Liliane
Hazra, Rishi
Ding, Shuangrui
Vaze, Sagar
Porcher, Francois
Li, Feng
Li, Siyuan
Kamath, Aishwarya
Cheng, Ho Kei
Dollár, Piotr
Ravi, Nikhila
Saenko, Kate
Zhang, Pengchuan
Feichtenhofer, Christoph
contents We present Segment Anything Model (SAM) 3, a unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short noun phrases (e.g., "yellow school bus"), image exemplars, or a combination of both. Promptable Concept Segmentation (PCS) takes such prompts and returns segmentation masks and unique identities for all matching object instances. To advance PCS, we build a scalable data engine that produces a high-quality dataset with 4M unique concept labels, including hard negatives, across images and videos. Our model consists of an image-level detector and a memory-based video tracker that share a single backbone. Recognition and localization are decoupled with a presence head, which boosts detection accuracy. SAM 3 doubles the accuracy of existing systems in both image and video PCS, and improves previous SAM capabilities on visual segmentation tasks. We open source SAM 3 along with our new Segment Anything with Concepts (SA-Co) benchmark for promptable concept segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16719
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAM 3: Segment Anything with Concepts
Carion, Nicolas
Gustafson, Laura
Hu, Yuan-Ting
Debnath, Shoubhik
Hu, Ronghang
Suris, Didac
Ryali, Chaitanya
Alwala, Kalyan Vasudev
Khedr, Haitham
Huang, Andrew
Lei, Jie
Ma, Tengyu
Guo, Baishan
Kalla, Arpit
Marks, Markus
Greer, Joseph
Wang, Meng
Sun, Peize
Rädle, Roman
Afouras, Triantafyllos
Mavroudi, Effrosyni
Xu, Katherine
Wu, Tsung-Han
Zhou, Yu
Momeni, Liliane
Hazra, Rishi
Ding, Shuangrui
Vaze, Sagar
Porcher, Francois
Li, Feng
Li, Siyuan
Kamath, Aishwarya
Cheng, Ho Kei
Dollár, Piotr
Ravi, Nikhila
Saenko, Kate
Zhang, Pengchuan
Feichtenhofer, Christoph
Computer Vision and Pattern Recognition
Artificial Intelligence
We present Segment Anything Model (SAM) 3, a unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short noun phrases (e.g., "yellow school bus"), image exemplars, or a combination of both. Promptable Concept Segmentation (PCS) takes such prompts and returns segmentation masks and unique identities for all matching object instances. To advance PCS, we build a scalable data engine that produces a high-quality dataset with 4M unique concept labels, including hard negatives, across images and videos. Our model consists of an image-level detector and a memory-based video tracker that share a single backbone. Recognition and localization are decoupled with a presence head, which boosts detection accuracy. SAM 3 doubles the accuracy of existing systems in both image and video PCS, and improves previous SAM capabilities on visual segmentation tasks. We open source SAM 3 along with our new Segment Anything with Concepts (SA-Co) benchmark for promptable concept segmentation.
title SAM 3: Segment Anything with Concepts
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.16719