Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Shuangrui, Qian, Rui, Xu, Haohang, Lin, Dahua, Xiong, Hongkai |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos
by: Qian, Rui, et al.
Published: (2023)
by: Qian, Rui, et al.
Published: (2023)
Rethinking Image-to-Video Adaptation: An Object-centric Perspective
by: Qian, Rui, et al.
Published: (2024)
by: Qian, Rui, et al.
Published: (2024)
Masked Autoencoders are Robust Data Augmentors
by: Xu, Haohang, et al.
Published: (2022)
by: Xu, Haohang, et al.
Published: (2022)
Streaming Long Video Understanding with Large Language Models
by: Qian, Rui, et al.
Published: (2024)
by: Qian, Rui, et al.
Published: (2024)
SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
by: Ding, Shuangrui, et al.
Published: (2024)
by: Ding, Shuangrui, et al.
Published: (2024)
Advancing Complex Video Object Segmentation via Progressive Concept Construction
by: Zhang, Zhixiong, et al.
Published: (2025)
by: Zhang, Zhixiong, et al.
Published: (2025)
MSF: Efficient Diffusion Model Via Multi-Scale Latent Factorize
by: Xu, Haohang, et al.
Published: (2025)
by: Xu, Haohang, et al.
Published: (2025)
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
by: Qian, Rui, et al.
Published: (2025)
by: Qian, Rui, et al.
Published: (2025)
Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention
by: Truong, Quang-Trung, et al.
Published: (2024)
by: Truong, Quang-Trung, et al.
Published: (2024)
A Simple yet Effective Network based on Vision Transformer for Camouflaged Object and Salient Object Detection
by: Hao, Chao, et al.
Published: (2024)
by: Hao, Chao, et al.
Published: (2024)
2nd Place Report of MOSEv2 Challenge 2025: Concept Guided Video Object Segmentation via SeC
by: Zhang, Zhixiong, et al.
Published: (2025)
by: Zhang, Zhixiong, et al.
Published: (2025)
A Simple yet Effective Subway Self-positioning Method based on Aerial-view Sleeper Detection
by: Song, Jiajie, et al.
Published: (2024)
by: Song, Jiajie, et al.
Published: (2024)
Image Compression for Machine and Human Vision with Spatial-Frequency Adaptation
by: Li, Han, et al.
Published: (2024)
by: Li, Han, et al.
Published: (2024)
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
S2-Track: A Simple yet Strong Approach for End-to-End 3D Multi-Object Tracking
by: Tang, Tao, et al.
Published: (2024)
by: Tang, Tao, et al.
Published: (2024)
Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image Editing
by: Lin, Haonan, et al.
Published: (2024)
by: Lin, Haonan, et al.
Published: (2024)
FIND: A Simple yet Effective Baseline for Diffusion-Generated Image Detection
by: Li, Jie, et al.
Published: (2026)
by: Li, Jie, et al.
Published: (2026)
A Simple yet Effective Test-Time Adaptation for Zero-Shot Monocular Metric Depth Estimation
by: Marsal, Rémi, et al.
Published: (2024)
by: Marsal, Rémi, et al.
Published: (2024)
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adapters
by: Lyu, Liangwei, et al.
Published: (2026)
by: Lyu, Liangwei, et al.
Published: (2026)
Guided Slot Attention for Unsupervised Video Object Segmentation
by: Lee, Minhyeok, et al.
Published: (2023)
by: Lee, Minhyeok, et al.
Published: (2023)
Dual Prototype Attention for Unsupervised Video Object Segmentation
by: Cho, Suhwan, et al.
Published: (2022)
by: Cho, Suhwan, et al.
Published: (2022)
HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse Poses
by: Ma, Caoyuan, et al.
Published: (2023)
by: Ma, Caoyuan, et al.
Published: (2023)
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
by: He, Ju, et al.
Published: (2023)
by: He, Ju, et al.
Published: (2023)
A Simple yet Effective Layout Token in Large Language Models for Document Understanding
by: Zhu, Zhaoqing, et al.
Published: (2025)
by: Zhu, Zhaoqing, et al.
Published: (2025)
CubeFormer: A Simple yet Effective Baseline for Lightweight Image Super-Resolution
by: Wang, Jikai, et al.
Published: (2024)
by: Wang, Jikai, et al.
Published: (2024)
Shifting Spotlight for Co-supervision: A Simple yet Efficient Single-branch Network to See Through Camouflage
by: Hu, Yang, et al.
Published: (2024)
by: Hu, Yang, et al.
Published: (2024)
Towards Diverse Binary Segmentation via A Simple yet General Gated Network
by: Zhao, Xiaoqi, et al.
Published: (2023)
by: Zhao, Xiaoqi, et al.
Published: (2023)
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
by: Kang, Seongjae, et al.
Published: (2025)
by: Kang, Seongjae, et al.
Published: (2025)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
by: Shaker, Abdelrahman, et al.
Published: (2024)
by: Shaker, Abdelrahman, et al.
Published: (2024)
Robust Video-Based Pothole Detection and Area Estimation for Intelligent Vehicles with Depth Map and Kalman Smoothing
by: Wang, Dehao, et al.
Published: (2025)
by: Wang, Dehao, et al.
Published: (2025)
Progressively Normalized Self-Attention Network for Video Polyp Segmentation
by: Ji, Ge-Peng, et al.
Published: (2021)
by: Ji, Ge-Peng, et al.
Published: (2021)
MOVE: Motion-Guided Few-Shot Video Object Segmentation
by: Ying, Kaining, et al.
Published: (2025)
by: Ying, Kaining, et al.
Published: (2025)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
by: Zhu, Chenming, et al.
Published: (2024)
by: Zhu, Chenming, et al.
Published: (2024)
LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
by: Yan, Cilin, et al.
Published: (2025)
by: Yan, Cilin, et al.
Published: (2025)
Slot-BERT: Self-supervised Object Discovery in Surgical Video
by: Liao, Guiqiu, et al.
Published: (2025)
by: Liao, Guiqiu, et al.
Published: (2025)
EventRR: Event Referential Reasoning for Referring Video Object Segmentation
by: Xu, Huihui, et al.
Published: (2025)
by: Xu, Huihui, et al.
Published: (2025)
Collaboratively Self-supervised Video Representation Learning for Action Recognition
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Self-supervised Video Instance Segmentation Can Boost Geographic Entity Alignment in Historical Maps
by: Xia, Xue, et al.
Published: (2024)
by: Xia, Xue, et al.
Published: (2024)
Event-assisted Low-Light Video Object Segmentation
by: Li, Hebei, et al.
Published: (2024)
by: Li, Hebei, et al.
Published: (2024)
Similar Items
-
Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos
by: Qian, Rui, et al.
Published: (2023) -
Rethinking Image-to-Video Adaptation: An Object-centric Perspective
by: Qian, Rui, et al.
Published: (2024) -
Masked Autoencoders are Robust Data Augmentors
by: Xu, Haohang, et al.
Published: (2022) -
Streaming Long Video Understanding with Large Language Models
by: Qian, Rui, et al.
Published: (2024) -
SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
by: Ding, Shuangrui, et al.
Published: (2024)