Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Haobo, Li, Xiangtai, Zhou, Chong, Li, Yining, Chen, Kai, Loy, Chen Change |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EdgeSAM: Prompt-In-the-Loop Distillation for SAM
by: Zhou, Chong, et al.
Published: (2023)
by: Zhou, Chong, et al.
Published: (2023)
OMG-Seg: Is One Model Good Enough For All Segmentation?
by: Li, Xiangtai, et al.
Published: (2024)
by: Li, Xiangtai, et al.
Published: (2024)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023)
by: Xu, Shilin, et al.
Published: (2023)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
by: Wu, Size, et al.
Published: (2023)
by: Wu, Size, et al.
Published: (2023)
Transformer-Based Visual Segmentation: A Survey
by: Li, Xiangtai, et al.
Published: (2023)
by: Li, Xiangtai, et al.
Published: (2023)
Mamba or RWKV: Exploring High-Quality and High-Efficiency Segment Anything Model
by: Yuan, Haobo, et al.
Published: (2024)
by: Yuan, Haobo, et al.
Published: (2024)
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
by: Xie, Jiahao, et al.
Published: (2023)
by: Xie, Jiahao, et al.
Published: (2023)
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
by: Xu, Shilin, et al.
Published: (2024)
by: Xu, Shilin, et al.
Published: (2024)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Control Color: Multimodal Diffusion-based Interactive Image Colorization
by: Liang, Zhexin, et al.
Published: (2024)
by: Liang, Zhexin, et al.
Published: (2024)
Explore In-Context Segmentation via Latent Diffusion Models
by: Wang, Chaoyang, et al.
Published: (2024)
by: Wang, Chaoyang, et al.
Published: (2024)
Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding
by: Tai, Hanchen, et al.
Published: (2024)
by: Tai, Hanchen, et al.
Published: (2024)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Kalman-Inspired Feature Propagation for Video Face Super-Resolution
by: Feng, Ruicheng, et al.
Published: (2024)
by: Feng, Ruicheng, et al.
Published: (2024)
Generalizable Implicit Motion Modeling for Video Frame Interpolation
by: Guo, Zujin, et al.
Published: (2024)
by: Guo, Zujin, et al.
Published: (2024)
AITTI: Learning Adaptive Inclusive Token for Text-to-Image Generation
by: Hou, Xinyu, et al.
Published: (2024)
by: Hou, Xinyu, et al.
Published: (2024)
Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
by: Li, Xiaoming, et al.
Published: (2025)
by: Li, Xiaoming, et al.
Published: (2025)
SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM
by: Chen, Lin, et al.
Published: (2025)
by: Chen, Lin, et al.
Published: (2025)
Trans-Adapter: A Plug-and-Play Framework for Transparent Image Inpainting
by: Dai, Yuekun, et al.
Published: (2025)
by: Dai, Yuekun, et al.
Published: (2025)
Effective SAM Combination for Open-Vocabulary Semantic Segmentation
by: Lee, Minhyeok, et al.
Published: (2024)
by: Lee, Minhyeok, et al.
Published: (2024)
Point-In-Context: Understanding Point Cloud via In-Context Learning
by: Liu, Mengyuan, et al.
Published: (2024)
by: Liu, Mengyuan, et al.
Published: (2024)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior
by: Chen, Honghua, et al.
Published: (2024)
by: Chen, Honghua, et al.
Published: (2024)
Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
by: Chen, Qi, et al.
Published: (2025)
by: Chen, Qi, et al.
Published: (2025)
Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
by: Li, Lin, et al.
Published: (2025)
by: Li, Lin, et al.
Published: (2025)
Eliminating Feature Ambiguity for Few-Shot Segmentation
by: Xu, Qianxiong, et al.
Published: (2024)
by: Xu, Qianxiong, et al.
Published: (2024)
Taking A Closer Look at Interacting Objects: Interaction-Aware Open Vocabulary Scene Graph Generation
by: Li, Lin, et al.
Published: (2025)
by: Li, Lin, et al.
Published: (2025)
Learning Inclusion Matching for Animation Paint Bucket Colorization
by: Dai, Yuekun, et al.
Published: (2024)
by: Dai, Yuekun, et al.
Published: (2024)
Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions
by: Siyao, Li, et al.
Published: (2025)
by: Siyao, Li, et al.
Published: (2025)
DifFace: Blind Face Restoration with Diffused Error Contraction
by: Yue, Zongsheng, et al.
Published: (2022)
by: Yue, Zongsheng, et al.
Published: (2022)
Omegance: A Single Parameter for Various Granularities in Diffusion-Based Synthesis
by: Hou, Xinyu, et al.
Published: (2024)
by: Hou, Xinyu, et al.
Published: (2024)
Towards Open Vocabulary Learning: A Survey
by: Wu, Jianzong, et al.
Published: (2023)
by: Wu, Jianzong, et al.
Published: (2023)
LA-Sign: Looped Transformers with Geometry-aware Alignment for Skeleton-based Sign Language Recognition
by: Pu, Muxin, et al.
Published: (2026)
by: Pu, Muxin, et al.
Published: (2026)
ObjCtrl-2.5D: Training-free Object Control with Camera Poses
by: Wang, Zhouxia, et al.
Published: (2024)
by: Wang, Zhouxia, et al.
Published: (2024)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
Learning 3D Garment Animation from Trajectories of A Piece of Cloth
by: Shao, Yidi, et al.
Published: (2025)
by: Shao, Yidi, et al.
Published: (2025)
Sigma: Semantically Informative Pre-training for Skeleton-based Sign Language Understanding
by: Pu, Muxin, et al.
Published: (2025)
by: Pu, Muxin, et al.
Published: (2025)
BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model
by: Song, Yiran, et al.
Published: (2024)
by: Song, Yiran, et al.
Published: (2024)
Taming SAM3 in the Wild: A Concept Bank for Open-Vocabulary Segmentation
by: Pei, Gensheng, et al.
Published: (2026)
by: Pei, Gensheng, et al.
Published: (2026)
CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation
by: Chen, Yanhui, et al.
Published: (2026)
by: Chen, Yanhui, et al.
Published: (2026)
Similar Items
-
EdgeSAM: Prompt-In-the-Loop Distillation for SAM
by: Zhou, Chong, et al.
Published: (2023) -
OMG-Seg: Is One Model Good Enough For All Segmentation?
by: Li, Xiangtai, et al.
Published: (2024) -
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023) -
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
by: Wu, Size, et al.
Published: (2023) -
Transformer-Based Visual Segmentation: A Survey
by: Li, Xiangtai, et al.
Published: (2023)