DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Mengshi, Zhu, Pengfei, Li, Xiangtai, Bi, Xiaoyang, Qi, Lu, Ma, Huadong, Yang, Ming-Hsuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Robust Unsupervised Attention Prediction in Autonomous Driving
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Uncovering the human motion pattern: Pattern Memory-based Diffusion Model for Trajectory Prediction
by: Yang, Yuxin, et al.
Published: (2024)
by: Yang, Yuxin, et al.
Published: (2024)
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
by: Xu, Shilin, et al.
Published: (2024)
by: Xu, Shilin, et al.
Published: (2024)
Multi-Stage Contrastive Regression for Action Quality Assessment
by: An, Qi, et al.
Published: (2024)
by: An, Qi, et al.
Published: (2024)
Question-Aware Evidence Ledgers for Video Relational Reasoning
by: Ou, Yilin, et al.
Published: (2026)
by: Ou, Yilin, et al.
Published: (2026)
Learning Group Interactions and Semantic Intentions for Multi-Object Trajectory Prediction
by: Qi, Mengshi, et al.
Published: (2024)
by: Qi, Mengshi, et al.
Published: (2024)
Mamba or RWKV: Exploring High-Quality and High-Efficiency Segment Anything Model
by: Yuan, Haobo, et al.
Published: (2024)
by: Yuan, Haobo, et al.
Published: (2024)
Robust Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Decomposed Vector-Quantized Variational Autoencoder for Human Grasp Generation
by: Zhao, Zhe, et al.
Published: (2024)
by: Zhao, Zhe, et al.
Published: (2024)
SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified Flow
by: Wang, Chaoyang, et al.
Published: (2024)
by: Wang, Chaoyang, et al.
Published: (2024)
BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model
by: Song, Yiran, et al.
Published: (2024)
by: Song, Yiran, et al.
Published: (2024)
Video Prediction Transformers without Recurrence or Convolution
by: Tang, Yujin, et al.
Published: (2024)
by: Tang, Yujin, et al.
Published: (2024)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
SAM3-I: Segment Anything with Instructions
by: Li, Jingjing, et al.
Published: (2025)
by: Li, Jingjing, et al.
Published: (2025)
Towards Efficient Object Re-Identification with A Novel Cloud-Edge Collaborative Framework
by: Wang, Chuanming, et al.
Published: (2024)
by: Wang, Chuanming, et al.
Published: (2024)
Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model
by: Huang, Kuan-Chih, et al.
Published: (2024)
by: Huang, Kuan-Chih, et al.
Published: (2024)
I-MedSAM: Implicit Medical Image Segmentation with Segment Anything
by: Wei, Xiaobao, et al.
Published: (2023)
by: Wei, Xiaobao, et al.
Published: (2023)
SAM 2: Segment Anything in Images and Videos
by: Ravi, Nikhila, et al.
Published: (2024)
by: Ravi, Nikhila, et al.
Published: (2024)
Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Medical SAM 2: Segment medical images as video via Segment Anything Model 2
by: Zhu, Jiayuan, et al.
Published: (2024)
by: Zhu, Jiayuan, et al.
Published: (2024)
SemiSAM: Enhancing Semi-Supervised Medical Image Segmentation via SAM-Assisted Consistency Regularization
by: Zhang, Yichi, et al.
Published: (2023)
by: Zhang, Yichi, et al.
Published: (2023)
Segment Anything Is Not Always Perfect: An Investigation of SAM on Different Real-world Applications
by: Ji, Wei, et al.
Published: (2023)
by: Ji, Wei, et al.
Published: (2023)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving
by: Lv, Changsheng, et al.
Published: (2024)
by: Lv, Changsheng, et al.
Published: (2024)
Improving Batch Normalization with TTA for Robust Object Detection in Self-Driving
by: Liao, Dacheng, et al.
Published: (2024)
by: Liao, Dacheng, et al.
Published: (2024)
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
by: Lin, Yuxin, et al.
Published: (2025)
by: Lin, Yuxin, et al.
Published: (2025)
Semi-Supervised Teacher-Reference-Student Architecture for Action Quality Assessment
by: Yun, Wulian, et al.
Published: (2024)
by: Yun, Wulian, et al.
Published: (2024)
A New Teacher-Reviewer-Student Framework for Semi-supervised 2D Human Pose Estimation
by: Yun, Wulian, et al.
Published: (2025)
by: Yun, Wulian, et al.
Published: (2025)
SAM-PD: How Far Can SAM Take Us in Tracking and Segmenting Anything in Videos by Prompt Denoising
by: Zhou, Tao, et al.
Published: (2024)
by: Zhou, Tao, et al.
Published: (2024)
Biomedical SAM 2: Segment Anything in Biomedical Images and Videos
by: Yan, Zhiling, et al.
Published: (2024)
by: Yan, Zhiling, et al.
Published: (2024)
AlignSAM: Aligning Segment Anything Model to Open Context via Reinforcement Learning
by: Huang, Duojun, et al.
Published: (2024)
by: Huang, Duojun, et al.
Published: (2024)
Unified Dense Prediction of Video Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
VIoTGPT: Learning to Schedule Vision Tools in LLMs towards Intelligent Video Internet of Things
by: Zhong, Yaoyao, et al.
Published: (2023)
by: Zhong, Yaoyao, et al.
Published: (2023)
Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
SAM3-Adapter: Efficient Adaptation of Segment Anything 3 for Camouflage Object Segmentation, Shadow Detection, and Medical Image Segmentation
by: Chen, Tianrun, et al.
Published: (2025)
by: Chen, Tianrun, et al.
Published: (2025)
MedSAM2: Segment Anything in 3D Medical Images and Videos
by: Ma, Jun, et al.
Published: (2025)
by: Ma, Jun, et al.
Published: (2025)
Mutual Distillation Learning For Person Re-Identification
by: Fu, Huiyuan, et al.
Published: (2024)
by: Fu, Huiyuan, et al.
Published: (2024)
CAT-SAM: Conditional Tuning for Few-Shot Adaptation of Segment Anything Model
by: Xiao, Aoran, et al.
Published: (2024)
by: Xiao, Aoran, et al.
Published: (2024)
SAM Struggles in Concealed Scenes -- Empirical Study on Segment Anything
by: Ji, Ge-Peng, et al.
Published: (2023)
by: Ji, Ge-Peng, et al.
Published: (2023)
Similar Items
-
Towards Robust Unsupervised Attention Prediction in Autonomous Driving
by: Qi, Mengshi, et al.
Published: (2025) -
Uncovering the human motion pattern: Pattern Memory-based Diffusion Model for Trajectory Prediction
by: Yang, Yuxin, et al.
Published: (2024) -
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
by: Xu, Shilin, et al.
Published: (2024) -
Multi-Stage Contrastive Regression for Action Quality Assessment
by: An, Qi, et al.
Published: (2024) -
Question-Aware Evidence Ledgers for Video Relational Reasoning
by: Ou, Yilin, et al.
Published: (2026)