OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Yeon-Ji, Kim, Jaein, Choi, Suhyung, Kim, Jin-Hwa, Zhang, Byoung-Tak |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DBMovi-GS: Dynamic View Synthesis from Blurry Monocular Video via Sparse-Controlled Gaussian Splatting
by: Song, Yeon-Ji, et al.
Published: (2025)
by: Song, Yeon-Ji, et al.
Published: (2025)
Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes
by: Song, Yeon-Ji, et al.
Published: (2026)
by: Song, Yeon-Ji, et al.
Published: (2026)
OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
by: Choi, Won-Seok, et al.
Published: (2025)
by: Choi, Won-Seok, et al.
Published: (2025)
Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis
by: Kim, Jaein, et al.
Published: (2026)
by: Kim, Jaein, et al.
Published: (2026)
Exploring Ordinal Bias in Action Recognition for Instructional Videos
by: Kim, Joochan, et al.
Published: (2025)
by: Kim, Joochan, et al.
Published: (2025)
Surface-Based Visibility-Guided Uncertainty for Continuous Active 3D Neural Reconstruction
by: Kim, Hyunseo, et al.
Published: (2024)
by: Kim, Hyunseo, et al.
Published: (2024)
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
by: Jung, Minjoon, et al.
Published: (2025)
by: Jung, Minjoon, et al.
Published: (2025)
Continual Vision-and-Language Navigation
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models
by: Lee, Hyundo, et al.
Published: (2025)
by: Lee, Hyundo, et al.
Published: (2025)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023)
by: Jung, Minjoon, et al.
Published: (2023)
PruNeRF: Segment-Centric Dataset Pruning via 3D Spatial Consistency
by: Jung, Yeonsung, et al.
Published: (2024)
by: Jung, Yeonsung, et al.
Published: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
by: Chi, Donghwan, et al.
Published: (2025)
by: Chi, Donghwan, et al.
Published: (2025)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025)
by: Jin, Hoiyeong, et al.
Published: (2025)
NBBOX: Noisy Bounding Box Improves Remote Sensing Object Detection
by: Kim, Yechan, et al.
Published: (2024)
by: Kim, Yechan, et al.
Published: (2024)
Voronoi-based Second-order Descriptor with Whitened Metric in LiDAR Place Recognition
by: Kim, Jaein, et al.
Published: (2026)
by: Kim, Jaein, et al.
Published: (2026)
Polyhedral Complex Derivation from Piecewise Trilinear Networks
by: Kim, Jin-Hwa
Published: (2024)
by: Kim, Jin-Hwa
Published: (2024)
Deformable Dynamic Convolution for Accurate yet Efficient Spatio-Temporal Traffic Prediction
by: Jin, Hyeonseok, et al.
Published: (2025)
by: Jin, Hyeonseok, et al.
Published: (2025)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
by: Gandhi, Sanket, et al.
Published: (2024)
by: Gandhi, Sanket, et al.
Published: (2024)
Instruction-tuned Self-Questioning Framework for Multimodal Reasoning
by: Jang, You-Won, et al.
Published: (2025)
by: Jang, You-Won, et al.
Published: (2025)
Synergistic Integration of Coordinate Network and Tensorial Feature for Improving Neural Radiance Fields from Sparse Inputs
by: Kim, Mingyu, et al.
Published: (2024)
by: Kim, Mingyu, et al.
Published: (2024)
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
by: Lee, Gayoung, et al.
Published: (2025)
by: Lee, Gayoung, et al.
Published: (2025)
Enhancing Self-Supervised Fine-Grained Video Object Tracking with Dynamic Memory Prediction
by: Zhou, Zihan, et al.
Published: (2025)
by: Zhou, Zihan, et al.
Published: (2025)
Object-Centric World Model for Language-Guided Manipulation
by: Jeong, Youngjoon, et al.
Published: (2025)
by: Jeong, Youngjoon, et al.
Published: (2025)
CogME: A Cognition-Inspired Multi-Dimensional Evaluation Metric for Story Understanding
by: Shin, Minjung, et al.
Published: (2021)
by: Shin, Minjung, et al.
Published: (2021)
Highly Efficient and Unsupervised Framework for Moving Object Detection in Satellite Videos
by: Xiao, C., et al.
Published: (2024)
by: Xiao, C., et al.
Published: (2024)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
by: Wang, Xucheng, et al.
Published: (2026)
by: Wang, Xucheng, et al.
Published: (2026)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
by: Zhou, Zhiyu, et al.
Published: (2026)
by: Zhou, Zhiyu, et al.
Published: (2026)
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
by: Li, Zhiyuan, et al.
Published: (2026)
by: Li, Zhiyuan, et al.
Published: (2026)
TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
by: Kim, Yoonsik, et al.
Published: (2024)
by: Kim, Yoonsik, et al.
Published: (2024)
Reasoning-Enhanced Object-Centric Learning for Videos
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object Segmentation
by: Xiangyu, Zheng, et al.
Published: (2025)
by: Xiangyu, Zheng, et al.
Published: (2025)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
PGA: Personalizing Grasping Agents with Single Human-Robot Interaction
by: Kim, Junghyun, et al.
Published: (2023)
by: Kim, Junghyun, et al.
Published: (2023)
What Happens When: Learning Temporal Orders of Events in Videos
by: Ahn, Daechul, et al.
Published: (2025)
by: Ahn, Daechul, et al.
Published: (2025)
Object-Centric Latent Action Learning
by: Klepach, Albina, et al.
Published: (2025)
by: Klepach, Albina, et al.
Published: (2025)
Exploiting Style Latent Flows for Generalizing Deepfake Video Detection
by: Choi, Jongwook, et al.
Published: (2024)
by: Choi, Jongwook, et al.
Published: (2024)
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
by: Zhang, Chenshuang, et al.
Published: (2025)
by: Zhang, Chenshuang, et al.
Published: (2025)
Similar Items
-
DBMovi-GS: Dynamic View Synthesis from Blurry Monocular Video via Sparse-Controlled Gaussian Splatting
by: Song, Yeon-Ji, et al.
Published: (2025) -
Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes
by: Song, Yeon-Ji, et al.
Published: (2026) -
OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
by: Choi, Won-Seok, et al.
Published: (2025) -
Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis
by: Kim, Jaein, et al.
Published: (2026) -
Exploring Ordinal Bias in Action Recognition for Instructional Videos
by: Kim, Joochan, et al.
Published: (2025)