OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
Fuente:
arXiv
Guardado en:
| Autores principales: | Song, Yeon-Ji, Kim, Jaein, Choi, Suhyung, Kim, Jin-Hwa, Zhang, Byoung-Tak |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DBMovi-GS: Dynamic View Synthesis from Blurry Monocular Video via Sparse-Controlled Gaussian Splatting
por: Song, Yeon-Ji, et al.
Publicado: (2025)
por: Song, Yeon-Ji, et al.
Publicado: (2025)
Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes
por: Song, Yeon-Ji, et al.
Publicado: (2026)
por: Song, Yeon-Ji, et al.
Publicado: (2026)
OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
por: Choi, Won-Seok, et al.
Publicado: (2025)
por: Choi, Won-Seok, et al.
Publicado: (2025)
Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis
por: Kim, Jaein, et al.
Publicado: (2026)
por: Kim, Jaein, et al.
Publicado: (2026)
Exploring Ordinal Bias in Action Recognition for Instructional Videos
por: Kim, Joochan, et al.
Publicado: (2025)
por: Kim, Joochan, et al.
Publicado: (2025)
Surface-Based Visibility-Guided Uncertainty for Continuous Active 3D Neural Reconstruction
por: Kim, Hyunseo, et al.
Publicado: (2024)
por: Kim, Hyunseo, et al.
Publicado: (2024)
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
por: Jung, Minjoon, et al.
Publicado: (2025)
por: Jung, Minjoon, et al.
Publicado: (2025)
Continual Vision-and-Language Navigation
por: Jeong, Seongjun, et al.
Publicado: (2024)
por: Jeong, Seongjun, et al.
Publicado: (2024)
Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models
por: Lee, Hyundo, et al.
Publicado: (2025)
por: Lee, Hyundo, et al.
Publicado: (2025)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
por: Grigore, Diana-Nicoleta, et al.
Publicado: (2025)
por: Grigore, Diana-Nicoleta, et al.
Publicado: (2025)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
por: Jeong, Seongjun, et al.
Publicado: (2024)
por: Jeong, Seongjun, et al.
Publicado: (2024)
OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
por: Kim, Juno, et al.
Publicado: (2025)
por: Kim, Juno, et al.
Publicado: (2025)
Background-aware Moment Detection for Video Moment Retrieval
por: Jung, Minjoon, et al.
Publicado: (2023)
por: Jung, Minjoon, et al.
Publicado: (2023)
PruNeRF: Segment-Centric Dataset Pruning via 3D Spatial Consistency
por: Jung, Yeonsung, et al.
Publicado: (2024)
por: Jung, Yeonsung, et al.
Publicado: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
por: Chi, Donghwan, et al.
Publicado: (2025)
por: Chi, Donghwan, et al.
Publicado: (2025)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
por: Jin, Hoiyeong, et al.
Publicado: (2025)
por: Jin, Hoiyeong, et al.
Publicado: (2025)
NBBOX: Noisy Bounding Box Improves Remote Sensing Object Detection
por: Kim, Yechan, et al.
Publicado: (2024)
por: Kim, Yechan, et al.
Publicado: (2024)
Voronoi-based Second-order Descriptor with Whitened Metric in LiDAR Place Recognition
por: Kim, Jaein, et al.
Publicado: (2026)
por: Kim, Jaein, et al.
Publicado: (2026)
Polyhedral Complex Derivation from Piecewise Trilinear Networks
por: Kim, Jin-Hwa
Publicado: (2024)
por: Kim, Jin-Hwa
Publicado: (2024)
Deformable Dynamic Convolution for Accurate yet Efficient Spatio-Temporal Traffic Prediction
por: Jin, Hyeonseok, et al.
Publicado: (2025)
por: Jin, Hyeonseok, et al.
Publicado: (2025)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
por: Gandhi, Sanket, et al.
Publicado: (2024)
por: Gandhi, Sanket, et al.
Publicado: (2024)
Instruction-tuned Self-Questioning Framework for Multimodal Reasoning
por: Jang, You-Won, et al.
Publicado: (2025)
por: Jang, You-Won, et al.
Publicado: (2025)
Synergistic Integration of Coordinate Network and Tensorial Feature for Improving Neural Radiance Fields from Sparse Inputs
por: Kim, Mingyu, et al.
Publicado: (2024)
por: Kim, Mingyu, et al.
Publicado: (2024)
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
por: Lee, Gayoung, et al.
Publicado: (2025)
por: Lee, Gayoung, et al.
Publicado: (2025)
Enhancing Self-Supervised Fine-Grained Video Object Tracking with Dynamic Memory Prediction
por: Zhou, Zihan, et al.
Publicado: (2025)
por: Zhou, Zihan, et al.
Publicado: (2025)
Object-Centric World Model for Language-Guided Manipulation
por: Jeong, Youngjoon, et al.
Publicado: (2025)
por: Jeong, Youngjoon, et al.
Publicado: (2025)
CogME: A Cognition-Inspired Multi-Dimensional Evaluation Metric for Story Understanding
por: Shin, Minjung, et al.
Publicado: (2021)
por: Shin, Minjung, et al.
Publicado: (2021)
Highly Efficient and Unsupervised Framework for Moving Object Detection in Satellite Videos
por: Xiao, C., et al.
Publicado: (2024)
por: Xiao, C., et al.
Publicado: (2024)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
por: Wang, Xucheng, et al.
Publicado: (2026)
por: Wang, Xucheng, et al.
Publicado: (2026)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
por: Zhou, Zhiyu, et al.
Publicado: (2026)
por: Zhou, Zhiyu, et al.
Publicado: (2026)
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
por: Li, Zhiyuan, et al.
Publicado: (2026)
por: Li, Zhiyuan, et al.
Publicado: (2026)
TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
por: Kim, Yoonsik, et al.
Publicado: (2024)
por: Kim, Yoonsik, et al.
Publicado: (2024)
Reasoning-Enhanced Object-Centric Learning for Videos
por: Li, Jian, et al.
Publicado: (2024)
por: Li, Jian, et al.
Publicado: (2024)
Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object Segmentation
por: Xiangyu, Zheng, et al.
Publicado: (2025)
por: Xiangyu, Zheng, et al.
Publicado: (2025)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
por: Cao, Tri, et al.
Publicado: (2026)
por: Cao, Tri, et al.
Publicado: (2026)
PGA: Personalizing Grasping Agents with Single Human-Robot Interaction
por: Kim, Junghyun, et al.
Publicado: (2023)
por: Kim, Junghyun, et al.
Publicado: (2023)
What Happens When: Learning Temporal Orders of Events in Videos
por: Ahn, Daechul, et al.
Publicado: (2025)
por: Ahn, Daechul, et al.
Publicado: (2025)
Object-Centric Latent Action Learning
por: Klepach, Albina, et al.
Publicado: (2025)
por: Klepach, Albina, et al.
Publicado: (2025)
Exploiting Style Latent Flows for Generalizing Deepfake Video Detection
por: Choi, Jongwook, et al.
Publicado: (2024)
por: Choi, Jongwook, et al.
Publicado: (2024)
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
por: Zhang, Chenshuang, et al.
Publicado: (2025)
por: Zhang, Chenshuang, et al.
Publicado: (2025)
Ejemplares similares
-
DBMovi-GS: Dynamic View Synthesis from Blurry Monocular Video via Sparse-Controlled Gaussian Splatting
por: Song, Yeon-Ji, et al.
Publicado: (2025) -
Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes
por: Song, Yeon-Ji, et al.
Publicado: (2026) -
OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
por: Choi, Won-Seok, et al.
Publicado: (2025) -
Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis
por: Kim, Jaein, et al.
Publicado: (2026) -
Exploring Ordinal Bias in Action Recognition for Instructional Videos
por: Kim, Joochan, et al.
Publicado: (2025)