InstructSAM: Segment Any Instance with Any Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Yuqian, Li, Wentong, Li, Zhaocheng, Lin, Yutong, Li, Juncheng, Tang, Siliang, Xiao, Jun, Zhuang, Yueting, Zhang, Wenqiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
by: Yu, Binhe, et al.
Published: (2025)
by: Yu, Binhe, et al.
Published: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition
by: Zheng, Yijie, et al.
Published: (2025)
by: Zheng, Yijie, et al.
Published: (2025)
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
by: Yu, Qifan, et al.
Published: (2024)
by: Yu, Qifan, et al.
Published: (2024)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
by: Miao, Bingchen, et al.
Published: (2024)
by: Miao, Bingchen, et al.
Published: (2024)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
Insight Any Instance: Promptable Instance Segmentation for Remote Sensing Images
by: Li, Xuexue
Published: (2024)
by: Li, Xuexue
Published: (2024)
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
by: Li, Juncheng, et al.
Published: (2023)
by: Li, Juncheng, et al.
Published: (2023)
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
by: Yuan, Yuqian, et al.
Published: (2025)
by: Yuan, Yuqian, et al.
Published: (2025)
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
by: Bu, Wendong, et al.
Published: (2025)
by: Bu, Wendong, et al.
Published: (2025)
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
by: Qiu, Haiyi, et al.
Published: (2026)
by: Qiu, Haiyi, et al.
Published: (2026)
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following
by: Li, Shufan, et al.
Published: (2023)
by: Li, Shufan, et al.
Published: (2023)
Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation
by: Zhang, Weiming, et al.
Published: (2026)
by: Zhang, Weiming, et al.
Published: (2026)
Compress Any Segment Anything Model (SAM)
by: Fan, Juntong, et al.
Published: (2025)
by: Fan, Juntong, et al.
Published: (2025)
X-SAM: From Segment Anything to Any Segmentation
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
by: Wang, Keyu, et al.
Published: (2026)
by: Wang, Keyu, et al.
Published: (2026)
Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
by: Yu, Qifan, et al.
Published: (2024)
by: Yu, Qifan, et al.
Published: (2024)
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
by: Zheng, Haoyu, et al.
Published: (2024)
by: Zheng, Haoyu, et al.
Published: (2024)
Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
by: Gao, Minghe, et al.
Published: (2024)
by: Gao, Minghe, et al.
Published: (2024)
X2SAM: Any Segmentation in Images and Videos
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
by: Gao, Minghe, et al.
Published: (2023)
by: Gao, Minghe, et al.
Published: (2023)
SAI3D: Segment Any Instance in 3D Scenes
by: Yin, Yingda, et al.
Published: (2023)
by: Yin, Yingda, et al.
Published: (2023)
Instruct2See: Learning to Remove Any Obstructions Across Distributions
by: Li, Junhang, et al.
Published: (2025)
by: Li, Junhang, et al.
Published: (2025)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
by: Ge, Zhiqi, et al.
Published: (2024)
by: Ge, Zhiqi, et al.
Published: (2024)
Segment Any Mesh
by: Tang, George, et al.
Published: (2024)
by: Tang, George, et al.
Published: (2024)
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
by: Zheng, Haoyu, et al.
Published: (2024)
by: Zheng, Haoyu, et al.
Published: (2024)
HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
by: Yu, Qifan, et al.
Published: (2023)
by: Yu, Qifan, et al.
Published: (2023)
Online Segment Any 3D Thing as Instance Tracking
by: Wang, Hanshi, et al.
Published: (2025)
by: Wang, Hanshi, et al.
Published: (2025)
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
MAS-SAM: Segment Any Marine Animal with Aggregated Features
by: Yan, Tianyu, et al.
Published: (2024)
by: Yan, Tianyu, et al.
Published: (2024)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
by: Fan, Zhaoyu, et al.
Published: (2025)
by: Fan, Zhaoyu, et al.
Published: (2025)
PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
by: Yuan, Yuqian, et al.
Published: (2025)
by: Yuan, Yuqian, et al.
Published: (2025)
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
by: Lin, Tianwei, et al.
Published: (2025)
by: Lin, Tianwei, et al.
Published: (2025)
SAM 2++: Tracking Anything at Any Granularity
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
Segment Any Medical Model Extended
by: Liu, Yihao, et al.
Published: (2024)
by: Liu, Yihao, et al.
Published: (2024)
Segment Any Change
by: Zheng, Zhuo, et al.
Published: (2024)
by: Zheng, Zhuo, et al.
Published: (2024)
EmbodiedSAM: Online Segment Any 3D Thing in Real Time
by: Xu, Xiuwei, et al.
Published: (2024)
by: Xu, Xiuwei, et al.
Published: (2024)
SAGOnline: Segment Any Gaussians Online
by: Sun, Wentao, et al.
Published: (2025)
by: Sun, Wentao, et al.
Published: (2025)
Similar Items
-
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
by: Yu, Binhe, et al.
Published: (2025) -
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023) -
InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition
by: Zheng, Yijie, et al.
Published: (2025) -
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
by: Yu, Qifan, et al.
Published: (2024) -
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
by: Miao, Bingchen, et al.
Published: (2024)