SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Ziyu, Zhang, Renrui, Zhu, Xiangyang, Tong, Chengzhuo, Gao, Peng, Li, Chunyuan, Heng, Pheng-Ann |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025)
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
von: Guo, Ziyu, et al.
Veröffentlicht: (2026)
von: Guo, Ziyu, et al.
Veröffentlicht: (2026)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
No Time to Train: Empowering Non-Parametric Networks for Few-shot 3D Scene Segmentation
von: Zhu, Xiangyang, et al.
Veröffentlicht: (2024)
von: Zhu, Xiangyang, et al.
Veröffentlicht: (2024)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Point-SAM: Promptable 3D Segmentation Model for Point Clouds
von: Zhou, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhou, Yuchen, et al.
Veröffentlicht: (2024)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
von: Li, Feng, et al.
Veröffentlicht: (2024)
von: Li, Feng, et al.
Veröffentlicht: (2024)
Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
Point Cloud Understanding via Attention-Driven Contrastive Learning
von: Wang, Yi, et al.
Veröffentlicht: (2024)
von: Wang, Yi, et al.
Veröffentlicht: (2024)
PointAD: Comprehending 3D Anomalies from Points and Pixels for Zero-shot 3D Anomaly Detection
von: Zhou, Qihang, et al.
Veröffentlicht: (2024)
von: Zhou, Qihang, et al.
Veröffentlicht: (2024)
Automated Safety Benchmarking: A Multi-agent Pipeline for LVLMs
von: Zhu, Xiangyang, et al.
Veröffentlicht: (2026)
von: Zhu, Xiangyang, et al.
Veröffentlicht: (2026)
Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning
von: Yang, Hui-Yue, et al.
Veröffentlicht: (2024)
von: Yang, Hui-Yue, et al.
Veröffentlicht: (2024)
Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding
von: Tang, Yiwen, et al.
Veröffentlicht: (2024)
von: Tang, Yiwen, et al.
Veröffentlicht: (2024)
Political DEBATE: Efficient Zero-shot and Few-shot Classifiers for Political Text
von: Burnham, Michael, et al.
Veröffentlicht: (2024)
von: Burnham, Michael, et al.
Veröffentlicht: (2024)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
Zero-Shot Refinement of Buildings' Segmentation Models using SAM
von: Mayladan, Ali, et al.
Veröffentlicht: (2023)
von: Mayladan, Ali, et al.
Veröffentlicht: (2023)
Unveiling the Generalization Power of Fine-Tuned Large Language Models
von: Yang, Haoran, et al.
Veröffentlicht: (2024)
von: Yang, Haoran, et al.
Veröffentlicht: (2024)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
Domain-Hierarchy Adaptation via Chain of Iterative Reasoning for Few-shot Hierarchical Text Classification
von: Ji, Ke, et al.
Veröffentlicht: (2024)
von: Ji, Ke, et al.
Veröffentlicht: (2024)
PARM: Pipeline-Adapted Reward Model
von: Fan, Xingyu, et al.
Veröffentlicht: (2026)
von: Fan, Xingyu, et al.
Veröffentlicht: (2026)
Learning to Describe for Predicting Zero-shot Drug-Drug Interactions
von: Zhu, Fangqi, et al.
Veröffentlicht: (2024)
von: Zhu, Fangqi, et al.
Veröffentlicht: (2024)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval
von: Li, Siting, et al.
Veröffentlicht: (2025)
von: Li, Siting, et al.
Veröffentlicht: (2025)
VideoSAM: Open-World Video Segmentation
von: Guo, Pinxue, et al.
Veröffentlicht: (2024)
von: Guo, Pinxue, et al.
Veröffentlicht: (2024)
QANA: LLM-based Question Generation and Network Analysis for Zero-shot Key Point Analysis and Beyond
von: Fukuma, Tomoki, et al.
Veröffentlicht: (2024)
von: Fukuma, Tomoki, et al.
Veröffentlicht: (2024)
SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
ProMISe: Promptable Medical Image Segmentation using SAM
von: Wang, Jinfeng, et al.
Veröffentlicht: (2024)
von: Wang, Jinfeng, et al.
Veröffentlicht: (2024)
T3: A Novel Zero-shot Transfer Learning Framework Iteratively Training on an Assistant Task for a Target Task
von: Tong, Xindi, et al.
Veröffentlicht: (2024)
von: Tong, Xindi, et al.
Veröffentlicht: (2024)
PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data
von: Zhu, Zhe, et al.
Veröffentlicht: (2025)
von: Zhu, Zhe, et al.
Veröffentlicht: (2025)
Overcoming Support Dilution for Robust Few-shot Semantic Segmentation
von: Tang, Wailing, et al.
Veröffentlicht: (2025)
von: Tang, Wailing, et al.
Veröffentlicht: (2025)
X2SAM: Any Segmentation in Images and Videos
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
MV-SAM: Multi-view Promptable Segmentation using Pointmap Guidance
von: Jeong, Yoonwoo, et al.
Veröffentlicht: (2026)
von: Jeong, Yoonwoo, et al.
Veröffentlicht: (2026)
CLIP-Guided SAM: Parameter-Efficient Semantic Conditioning for Promptable Segmentation
von: Jalilian, Shayan, et al.
Veröffentlicht: (2026)
von: Jalilian, Shayan, et al.
Veröffentlicht: (2026)
Insight Any Instance: Promptable Instance Segmentation for Remote Sensing Images
von: Li, Xuexue
Veröffentlicht: (2024)
von: Li, Xuexue
Veröffentlicht: (2024)
DeID-GPT: Zero-shot Medical Text De-Identification by GPT-4
von: Liu, Zhengliang, et al.
Veröffentlicht: (2023)
von: Liu, Zhengliang, et al.
Veröffentlicht: (2023)
Robust Promptable Video Object Segmentation
von: Lee, Sohyun, et al.
Veröffentlicht: (2026)
von: Lee, Sohyun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025) -
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
von: Guo, Ziyu, et al.
Veröffentlicht: (2025) -
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
von: Guo, Ziyu, et al.
Veröffentlicht: (2025) -
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
von: Guo, Ziyu, et al.
Veröffentlicht: (2026) -
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)