Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qian, Rui, Ding, Shuangrui, Liu, Xian, Lin, Dahua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking Image-to-Video Adaptation: An Object-centric Perspective
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2023)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2023)
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
von: Qian, Rui, et al.
Veröffentlicht: (2025)
von: Qian, Rui, et al.
Veröffentlicht: (2025)
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024)
Advancing Complex Video Object Segmentation via Progressive Concept Construction
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025)
Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention
von: Truong, Quang-Trung, et al.
Veröffentlicht: (2024)
von: Truong, Quang-Trung, et al.
Veröffentlicht: (2024)
2nd Place Report of MOSEv2 Challenge 2025: Concept Guided Video Object Segmentation via SeC
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025)
Independently Keypoint Learning for Small Object Semantic Correspondence
von: Jin, Hailong, et al.
Veröffentlicht: (2024)
von: Jin, Hailong, et al.
Veröffentlicht: (2024)
Self-supervised Shape Completion via Involution and Implicit Correspondences
von: Liu, Mengya, et al.
Veröffentlicht: (2024)
von: Liu, Mengya, et al.
Veröffentlicht: (2024)
Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
Dynamic in Static: Hybrid Visual Correspondence for Self-Supervised Video Object Segmentation
von: Pei, Gensheng, et al.
Veröffentlicht: (2024)
von: Pei, Gensheng, et al.
Veröffentlicht: (2024)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
von: Xu, Jilan, et al.
Veröffentlicht: (2025)
von: Xu, Jilan, et al.
Veröffentlicht: (2025)
Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
von: Dong, Yuhao, et al.
Veröffentlicht: (2026)
von: Dong, Yuhao, et al.
Veröffentlicht: (2026)
Collaboratively Self-supervised Video Representation Learning for Action Recognition
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
Walker: Self-supervised Multiple Object Tracking by Walking on Temporal Appearance Graphs
von: Segu, Mattia, et al.
Veröffentlicht: (2024)
von: Segu, Mattia, et al.
Veröffentlicht: (2024)
Object-centric Video Question Answering with Visual Grounding and Referring
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
von: Fu, Xiao, et al.
Veröffentlicht: (2025)
von: Fu, Xiao, et al.
Veröffentlicht: (2025)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
Joint Spatial-Temporal Modeling and Contrastive Learning for Self-supervised Heart Rate Measurement
von: Qian, Wei, et al.
Veröffentlicht: (2024)
von: Qian, Wei, et al.
Veröffentlicht: (2024)
Learning Correspondence for Deformable Objects
von: Sundaresan, Priya, et al.
Veröffentlicht: (2024)
von: Sundaresan, Priya, et al.
Veröffentlicht: (2024)
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
von: Li, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Li, Zhiyuan, et al.
Veröffentlicht: (2026)
Plenoptic Video Generation
von: Fu, Xiao, et al.
Veröffentlicht: (2026)
von: Fu, Xiao, et al.
Veröffentlicht: (2026)
Emergent Temporal Correspondences from Video Diffusion Transformers
von: Nam, Jisu, et al.
Veröffentlicht: (2025)
von: Nam, Jisu, et al.
Veröffentlicht: (2025)
PicoPose: Progressive Pixel-to-Pixel Correspondence Learning for Novel Object Pose Estimation
von: Liu, Lihua, et al.
Veröffentlicht: (2025)
von: Liu, Lihua, et al.
Veröffentlicht: (2025)
Image Re-Identification: Where Self-supervision Meets Vision-Language Learning
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
von: Dünkel, Olaf, et al.
Veröffentlicht: (2026)
von: Dünkel, Olaf, et al.
Veröffentlicht: (2026)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
von: Sun, Guolei, et al.
Veröffentlicht: (2022)
von: Sun, Guolei, et al.
Veröffentlicht: (2022)
Leveraging Motion Information for Better Self-Supervised Video Correspondence Learning
von: Zhou, Zihan, et al.
Veröffentlicht: (2025)
von: Zhou, Zihan, et al.
Veröffentlicht: (2025)
DO3D: Self-supervised Learning of Decomposed Object-aware 3D Motion and Depth from Monocular Videos
von: Wu, Xiuzhe, et al.
Veröffentlicht: (2024)
von: Wu, Xiuzhe, et al.
Veröffentlicht: (2024)
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
von: Sun, Xiaokun, et al.
Veröffentlicht: (2026)
von: Sun, Xiaokun, et al.
Veröffentlicht: (2026)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
von: Yuan, Tianyu, et al.
Veröffentlicht: (2025)
von: Yuan, Tianyu, et al.
Veröffentlicht: (2025)
Masked Autoencoders are Robust Data Augmentors
von: Xu, Haohang, et al.
Veröffentlicht: (2022)
von: Xu, Haohang, et al.
Veröffentlicht: (2022)
Multi-granularity Correspondence Learning from Long-term Noisy Videos
von: Lin, Yijie, et al.
Veröffentlicht: (2024)
von: Lin, Yijie, et al.
Veröffentlicht: (2024)
SIRST-5K: Exploring Massive Negatives Synthesis with Self-supervised Learning for Robust Infrared Small Target Detection
von: Lu, Yahao, et al.
Veröffentlicht: (2024)
von: Lu, Yahao, et al.
Veröffentlicht: (2024)
One-Shot Learning Meets Depth Diffusion in Multi-Object Videos
von: Jain, Anisha
Veröffentlicht: (2024)
von: Jain, Anisha
Veröffentlicht: (2024)
Deep Learning on Object-centric 3D Neural Fields
von: Ramirez, Pierluigi Zama, et al.
Veröffentlicht: (2023)
von: Ramirez, Pierluigi Zama, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Rethinking Image-to-Video Adaptation: An Object-centric Perspective
von: Qian, Rui, et al.
Veröffentlicht: (2024) -
Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2023) -
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024) -
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
von: Qian, Rui, et al.
Veröffentlicht: (2025) -
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
von: Liu, Yang, et al.
Veröffentlicht: (2025)