Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhiyuan, Zhao, Rongzhen, Yang, Wenyan, Zhao, Wenshuai, Marttinen, Pekka, Pajarinen, Joni |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Internalizing Temporal Consistency in Video Object-Centric Learning without Explicit Regularization
by: Zhao, Rongzhen, et al.
Published: (2026)
by: Zhao, Rongzhen, et al.
Published: (2026)
Cycle Consistency in Video Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2026)
by: Zhao, Rongzhen, et al.
Published: (2026)
Object-Centric Vision Token Pruning for Vision Language Models
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
Organized Grouped Discrete Representation for Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2024)
by: Zhao, Rongzhen, et al.
Published: (2024)
Vector-Quantized Vision Foundation Models for Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
Grouped Discrete Representation for Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2024)
by: Zhao, Rongzhen, et al.
Published: (2024)
Grouped Discrete Representation Guides Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2024)
by: Zhao, Rongzhen, et al.
Published: (2024)
MetaSlot: Break Through the Fixed Number of Slots in Object-Centric Learning
by: Liu, Hongjia, et al.
Published: (2025)
by: Liu, Hongjia, et al.
Published: (2025)
Smoothing Slot Attention Iterations and Recurrences
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
Predicting Video Slot Attention Queries from Random Slot-Feature Pairs
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
Multi-Scale Fusion for Object Representation
by: Zhao, Rongzhen, et al.
Published: (2024)
by: Zhao, Rongzhen, et al.
Published: (2024)
Slot Attention with Re-Initialization and Self-Distillation
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
Learning Progress Driven Multi-Agent Curriculum
by: Zhao, Wenshuai, et al.
Published: (2022)
by: Zhao, Wenshuai, et al.
Published: (2022)
Temporally Consistent Object-Centric Learning by Contrasting Slots
by: Manasyan, Anna, et al.
Published: (2024)
by: Manasyan, Anna, et al.
Published: (2024)
Closed-Loop Vision-Language Planning for Multi-Agent Coordination
by: Li, Zhiyuan, et al.
Published: (2025)
by: Li, Zhiyuan, et al.
Published: (2025)
RGB-Th-Bench: A Dense benchmark for Visual-Thermal Understanding of Vision Language Models
by: Moshtaghi, Mehdi, et al.
Published: (2025)
by: Moshtaghi, Mehdi, et al.
Published: (2025)
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
by: Song, Yeon-Ji, et al.
Published: (2024)
by: Song, Yeon-Ji, et al.
Published: (2024)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
by: Zou, Zhengtao, et al.
Published: (2025)
by: Zou, Zhengtao, et al.
Published: (2025)
Reasoning-Enhanced Object-Centric Learning for Videos
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing
by: Li, Zhiyuan, et al.
Published: (2026)
by: Li, Zhiyuan, et al.
Published: (2026)
Latent-Compressed Variational Autoencoder for Video Diffusion Models
by: Guan, Jiarui, et al.
Published: (2026)
by: Guan, Jiarui, et al.
Published: (2026)
Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
FastInit: Fast Noise Initialization for Temporally Consistent Video Generation
by: Bai, Chengyu, et al.
Published: (2025)
by: Bai, Chengyu, et al.
Published: (2025)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
by: Gao, Mingze, et al.
Published: (2024)
by: Gao, Mingze, et al.
Published: (2024)
AgentMixer: Multi-Agent Correlated Policy Factorization
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
Object-Centric Latent Action Learning
by: Klepach, Albina, et al.
Published: (2025)
by: Klepach, Albina, et al.
Published: (2025)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
by: Zhang, Runze, et al.
Published: (2025)
by: Zhang, Runze, et al.
Published: (2025)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
by: Luo, Yuanhao, et al.
Published: (2026)
by: Luo, Yuanhao, et al.
Published: (2026)
CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style Transfer
by: Nie, Wenbo, et al.
Published: (2026)
by: Nie, Wenbo, et al.
Published: (2026)
Strategies for Robust Deep Learning Based Deformable Registration
by: Honkamaa, Joel, et al.
Published: (2025)
by: Honkamaa, Joel, et al.
Published: (2025)
PAWS: Perception of Articulation in the Wild at Scale from Egocentric Videos
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
by: Gandhi, Sanket, et al.
Published: (2024)
by: Gandhi, Sanket, et al.
Published: (2024)
Revisiting Salient Object Detection from an Observer-Centric Perspective
by: Zhang, Fuxi, et al.
Published: (2026)
by: Zhang, Fuxi, et al.
Published: (2026)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
Hierarchical Object-Centric Learning with Capsule Networks
by: Renzulli, Riccardo
Published: (2024)
by: Renzulli, Riccardo
Published: (2024)
From Frames to Sequences: Temporally Consistent Human-Centric Dense Prediction
by: Miao, Xingyu, et al.
Published: (2026)
by: Miao, Xingyu, et al.
Published: (2026)
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
by: Li, Zhengqin, et al.
Published: (2026)
by: Li, Zhengqin, et al.
Published: (2026)
UVCG: Leveraging Temporal Consistency for Universal Video Protection
by: Li, KaiZhou, et al.
Published: (2024)
by: Li, KaiZhou, et al.
Published: (2024)
Similar Items
-
Internalizing Temporal Consistency in Video Object-Centric Learning without Explicit Regularization
by: Zhao, Rongzhen, et al.
Published: (2026) -
Cycle Consistency in Video Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2026) -
Object-Centric Vision Token Pruning for Vision Language Models
by: Li, Guangyuan, et al.
Published: (2025) -
Organized Grouped Discrete Representation for Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2024) -
Vector-Quantized Vision Foundation Models for Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2025)