Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tan, Yuting, Cheng, Xilong, Qin, Yunxiao, Li, Zhengnan, Zhang, Jingjing |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Leveraging Transformers for Weakly Supervised Object Localization in Unconstrained Videos
par: Murtaza, Shakeeb, et autres
Publié: (2024)
par: Murtaza, Shakeeb, et autres
Publié: (2024)
Towards Unconstrained Human-Object Interaction
par: Tonini, Francesco, et autres
Publié: (2026)
par: Tonini, Francesco, et autres
Publié: (2026)
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
par: Xu, Boshen, et autres
Publié: (2025)
par: Xu, Boshen, et autres
Publié: (2025)
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos
par: Yuan, Chengbo, et autres
Publié: (2024)
par: Yuan, Chengbo, et autres
Publié: (2024)
Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos
par: Zhu, Zhifan, et autres
Publié: (2023)
par: Zhu, Zhifan, et autres
Publié: (2023)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
par: Pei, Baoqi, et autres
Publié: (2025)
par: Pei, Baoqi, et autres
Publié: (2025)
Anticipating Next Active Objects for Egocentric Videos
par: Thakur, Sanket, et autres
Publié: (2023)
par: Thakur, Sanket, et autres
Publié: (2023)
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
par: Xu, Boshen, et autres
Publié: (2024)
par: Xu, Boshen, et autres
Publié: (2024)
WildActor: Unconstrained Identity-Preserving Video Generation
par: Guo, Qin, et autres
Publié: (2026)
par: Guo, Qin, et autres
Publié: (2026)
FTMixer: Frequency and Time Domain Representations Fusion for Time Series Modeling
par: Li, Zhengnan, et autres
Publié: (2024)
par: Li, Zhengnan, et autres
Publié: (2024)
Distillation-guided Representation Learning for Unconstrained Gait Recognition
par: Guo, Yuxiang, et autres
Publié: (2023)
par: Guo, Yuxiang, et autres
Publié: (2023)
WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos
par: Ye, Yufei, et autres
Publié: (2026)
par: Ye, Yufei, et autres
Publié: (2026)
Temporally Grounding Instructional Diagrams in Unconstrained Videos
par: Zhang, Jiahao, et autres
Publié: (2024)
par: Zhang, Jiahao, et autres
Publié: (2024)
Rethinking Detecting Salient and Camouflaged Objects in Unconstrained Scenes
par: Zhou, Zhangjun, et autres
Publié: (2024)
par: Zhou, Zhangjun, et autres
Publié: (2024)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
par: Ren, Sucheng, et autres
Publié: (2024)
par: Ren, Sucheng, et autres
Publié: (2024)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
par: Zhu, Zhifan, et autres
Publié: (2025)
par: Zhu, Zhifan, et autres
Publié: (2025)
VideoSSR: Video Self-Supervised Reinforcement Learning
par: He, Zefeng, et autres
Publié: (2025)
par: He, Zefeng, et autres
Publié: (2025)
Object-Shot Enhanced Grounding Network for Egocentric Video
par: Feng, Yisen, et autres
Publié: (2025)
par: Feng, Yisen, et autres
Publié: (2025)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
par: Xu, Zhiyang, et autres
Publié: (2026)
par: Xu, Zhiyang, et autres
Publié: (2026)
Semi-Supervised Unconstrained Head Pose Estimation in the Wild
par: Zhou, Huayi, et autres
Publié: (2024)
par: Zhou, Huayi, et autres
Publié: (2024)
From Understanding to Erasing: Towards Complete and Stable Video Object Removal
par: Liu, Dingming, et autres
Publié: (2026)
par: Liu, Dingming, et autres
Publié: (2026)
Enhancing Representations through Heterogeneous Self-Supervised Learning
par: Li, Zhong-Yu, et autres
Publié: (2023)
par: Li, Zhong-Yu, et autres
Publié: (2023)
Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
par: Dou, Zi-Yi, et autres
Publié: (2024)
par: Dou, Zi-Yi, et autres
Publié: (2024)
PersonaAnimator: Personalized Motion Transfer from Unconstrained Videos
par: Qian, Ziyun, et autres
Publié: (2025)
par: Qian, Ziyun, et autres
Publié: (2025)
FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video
par: Camiletto, Andrea Boscolo, et autres
Publié: (2025)
par: Camiletto, Andrea Boscolo, et autres
Publié: (2025)
Memory Storyboard: Leveraging Temporal Segmentation for Streaming Self-Supervised Learning from Egocentric Videos
par: Yang, Yanlai, et autres
Publié: (2025)
par: Yang, Yanlai, et autres
Publié: (2025)
CoLo-CAM: Class Activation Mapping for Object Co-Localization in Weakly-Labeled Unconstrained Videos
par: Belharbi, Soufiane, et autres
Publié: (2023)
par: Belharbi, Soufiane, et autres
Publié: (2023)
Self-Classification Enhancement and Correction for Weakly Supervised Object Detection
par: Yin, Yufei, et autres
Publié: (2025)
par: Yin, Yufei, et autres
Publié: (2025)
Self-Supervised Point Cloud Completion based on Multi-View Augmentations of Single Partial Point Cloud
par: Lu, Jingjing, et autres
Publié: (2025)
par: Lu, Jingjing, et autres
Publié: (2025)
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
par: Ahamed, Shihab Aaqil, et autres
Publié: (2025)
par: Ahamed, Shihab Aaqil, et autres
Publié: (2025)
A Generative Framework for Self-Supervised Facial Representation Learning
par: He, Ruian, et autres
Publié: (2023)
par: He, Ruian, et autres
Publié: (2023)
Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation
par: Zheng, Xiaoyun, et autres
Publié: (2025)
par: Zheng, Xiaoyun, et autres
Publié: (2025)
Self-Supervised Video Representation Learning in a Heuristic Decoupled Perspective
par: Song, Zeen, et autres
Publié: (2024)
par: Song, Zeen, et autres
Publié: (2024)
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
par: Su, Yuejiao, et autres
Publié: (2026)
par: Su, Yuejiao, et autres
Publié: (2026)
Dynamic in Static: Hybrid Visual Correspondence for Self-Supervised Video Object Segmentation
par: Pei, Gensheng, et autres
Publié: (2024)
par: Pei, Gensheng, et autres
Publié: (2024)
Towards Continual Egocentric Activity Recognition: A Multi-modal Egocentric Activity Dataset for Continual Learning
par: Xu, Linfeng, et autres
Publié: (2023)
par: Xu, Linfeng, et autres
Publié: (2023)
Retrieval-Augmented Egocentric Video Captioning
par: Xu, Jilan, et autres
Publié: (2024)
par: Xu, Jilan, et autres
Publié: (2024)
Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention
par: Liu, Haijing, et autres
Publié: (2025)
par: Liu, Haijing, et autres
Publié: (2025)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
par: Kumar, Akash, et autres
Publié: (2025)
par: Kumar, Akash, et autres
Publié: (2025)
Gaussian in the Wild: 3D Gaussian Splatting for Unconstrained Image Collections
par: Zhang, Dongbin, et autres
Publié: (2024)
par: Zhang, Dongbin, et autres
Publié: (2024)
Documents similaires
-
Leveraging Transformers for Weakly Supervised Object Localization in Unconstrained Videos
par: Murtaza, Shakeeb, et autres
Publié: (2024) -
Towards Unconstrained Human-Object Interaction
par: Tonini, Francesco, et autres
Publié: (2026) -
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
par: Xu, Boshen, et autres
Publié: (2025) -
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos
par: Yuan, Chengbo, et autres
Publié: (2024) -
Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos
par: Zhu, Zhifan, et autres
Publié: (2023)