Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Klepachevskyi, Dmytro, Wong, Alexander, Rambhatla, Sirisha, Chen, Yuhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAMJAM: Zero-Shot Video Scene Graph Generation for Egocentric Kitchen Videos
by: Li, Joshua, et al.
Published: (2025)
by: Li, Joshua, et al.
Published: (2025)
Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation
by: Bright, Jerrin, et al.
Published: (2025)
by: Bright, Jerrin, et al.
Published: (2025)
Domain-Guided Masked Autoencoders for Unique Player Identification
by: Balaji, Bavesh, et al.
Published: (2024)
by: Balaji, Bavesh, et al.
Published: (2024)
LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image Editing
by: Soni, Achint, et al.
Published: (2025)
by: Soni, Achint, et al.
Published: (2025)
DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
by: Delatolas, Thanos, et al.
Published: (2025)
by: Delatolas, Thanos, et al.
Published: (2025)
KitchenTwin: Semantically and Geometrically Grounded 3D Kitchen Digital Twins
by: Wu, Quanyun, et al.
Published: (2026)
by: Wu, Quanyun, et al.
Published: (2026)
Object-Shot Enhanced Grounding Network for Egocentric Video
by: Feng, Yisen, et al.
Published: (2025)
by: Feng, Yisen, et al.
Published: (2025)
Zero-Shot Temporal Interaction Localization for Egocentric Videos
by: Zhang, Erhang, et al.
Published: (2025)
by: Zhang, Erhang, et al.
Published: (2025)
Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification
by: Liu, Yangyang, et al.
Published: (2025)
by: Liu, Yangyang, et al.
Published: (2025)
RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM
by: Song, Ziying, et al.
Published: (2024)
by: Song, Ziying, et al.
Published: (2024)
SAM3D: Zero-Shot 3D Object Detection via Segment Anything Model
by: Zhang, Dingyuan, et al.
Published: (2023)
by: Zhang, Dingyuan, et al.
Published: (2023)
IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
FoodTrack: Estimating Handheld Food Portions with Egocentric Video
by: Wang, Ervin, et al.
Published: (2025)
by: Wang, Ervin, et al.
Published: (2025)
PinPoint: Prompting with Informative Interior Points
by: Sadeghi, Pouya, et al.
Published: (2026)
by: Sadeghi, Pouya, et al.
Published: (2026)
Zero Shot Context-Based Object Segmentation using SLIP (SAM+CLIP)
by: Gundavarapu, Saaketh Koundinya, et al.
Published: (2024)
by: Gundavarapu, Saaketh Koundinya, et al.
Published: (2024)
Zero-Shot Multi-Object Scene Completion
by: Iwase, Shun, et al.
Published: (2024)
by: Iwase, Shun, et al.
Published: (2024)
SAM-Mamba: Mamba Guided SAM Architecture for Generalized Zero-Shot Polyp Segmentation
by: Dutta, Tapas Kumar, et al.
Published: (2024)
by: Dutta, Tapas Kumar, et al.
Published: (2024)
Identification of Conversation Partners from Egocentric Video
by: Dorszewski, Tobias, et al.
Published: (2024)
by: Dorszewski, Tobias, et al.
Published: (2024)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
by: Li, Yingxuan, et al.
Published: (2024)
by: Li, Yingxuan, et al.
Published: (2024)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery
by: Sethi, Geet, et al.
Published: (2026)
by: Sethi, Geet, et al.
Published: (2026)
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
ClipSAM: CLIP and SAM Collaboration for Zero-Shot Anomaly Segmentation
by: Li, Shengze, et al.
Published: (2024)
by: Li, Shengze, et al.
Published: (2024)
SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation
by: Lin, Jiehong, et al.
Published: (2023)
by: Lin, Jiehong, et al.
Published: (2023)
Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
by: Yang, Zaiquan, et al.
Published: (2025)
by: Yang, Zaiquan, et al.
Published: (2025)
Re-Prompting SAM 3 via Object Retrieval: 3rd of the 5th PVUW MOSE Track
by: Gao, Mingqi, et al.
Published: (2026)
by: Gao, Mingqi, et al.
Published: (2026)
Comparing SAM 2 and SAM 3 for Zero-Shot Segmentation of 3D Medical Data
by: Chakrabarty, Satrajit, et al.
Published: (2025)
by: Chakrabarty, Satrajit, et al.
Published: (2025)
Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos
by: Tschernezki, Vadim, et al.
Published: (2025)
by: Tschernezki, Vadim, et al.
Published: (2025)
SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation
by: Xu, Mutian, et al.
Published: (2023)
by: Xu, Mutian, et al.
Published: (2023)
MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
by: Huang, Binhua, et al.
Published: (2025)
by: Huang, Binhua, et al.
Published: (2025)
Multi-Stage VLM Pipeline for Zero-Shot Traffic Accident Understanding
by: Tatematsu, Fumiya, et al.
Published: (2026)
by: Tatematsu, Fumiya, et al.
Published: (2026)
STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification
by: Xu, Xingguo, et al.
Published: (2026)
by: Xu, Xingguo, et al.
Published: (2026)
Zero-Shot Personalization of Objects via Textual Inversion
by: Roy, Aniket, et al.
Published: (2026)
by: Roy, Aniket, et al.
Published: (2026)
Motion-Zero: Zero-Shot Moving Object Control Framework for Diffusion-Based Video Generation
by: Chen, Changgu, et al.
Published: (2024)
by: Chen, Changgu, et al.
Published: (2024)
Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM
by: Wan, Jia, et al.
Published: (2024)
by: Wan, Jia, et al.
Published: (2024)
Universal Features Guided Zero-Shot Category-Level Object Pose Estimation
by: Qu, Wentian, et al.
Published: (2025)
by: Qu, Wentian, et al.
Published: (2025)
AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2026)
by: Jin, Woojeong, et al.
Published: (2026)
Zero-Shot Monocular Motion Segmentation in the Wild by Combining Deep Learning with Geometric Motion Model Fusion
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
Similar Items
-
SAMJAM: Zero-Shot Video Scene Graph Generation for Egocentric Kitchen Videos
by: Li, Joshua, et al.
Published: (2025) -
Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation
by: Bright, Jerrin, et al.
Published: (2025) -
Domain-Guided Masked Autoencoders for Unique Player Identification
by: Balaji, Bavesh, et al.
Published: (2024) -
LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image Editing
by: Soni, Achint, et al.
Published: (2025) -
DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification
by: Wang, Yuhao, et al.
Published: (2024)