Learning Object Permanence from Videos via Latent Imaginations
Fuente:
arXiv
Guardado en:
| Autores principales: | Traub, Manuel, Becker, Frederic, Otte, Sebastian, Butz, Martin V. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Loci-Segmented: Improving Scene Segmentation Learning
por: Traub, Manuel, et al.
Publicado: (2023)
por: Traub, Manuel, et al.
Publicado: (2023)
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
por: Traub, Manuel, et al.
Publicado: (2025)
por: Traub, Manuel, et al.
Publicado: (2025)
Offline Tracking with Object Permanence
por: Liu, Xianzhong, et al.
Publicado: (2023)
por: Liu, Xianzhong, et al.
Publicado: (2023)
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
por: Chen, Jiahe, et al.
Publicado: (2026)
por: Chen, Jiahe, et al.
Publicado: (2026)
Data-driven Verification of DNNs for Object Recognition
por: Otte, Clemens, et al.
Publicado: (2024)
por: Otte, Clemens, et al.
Publicado: (2024)
Detection of Fast-Moving Objects with Neuromorphic Hardware
por: Ziegler, Andreas, et al.
Publicado: (2024)
por: Ziegler, Andreas, et al.
Publicado: (2024)
Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation
por: Lou, Zijie, et al.
Publicado: (2026)
por: Lou, Zijie, et al.
Publicado: (2026)
Imagining the Unseen: Generative Location Modeling for Object Placement
por: Yun, Jooyeol, et al.
Publicado: (2024)
por: Yun, Jooyeol, et al.
Publicado: (2024)
Jigsaw++: Imagining Complete Shape Priors for Object Reassembly
por: Lu, Jiaxin, et al.
Publicado: (2024)
por: Lu, Jiaxin, et al.
Publicado: (2024)
Imagine360: Immersive 360 Video Generation from Perspective Anchor
por: Tan, Jing, et al.
Publicado: (2024)
por: Tan, Jing, et al.
Publicado: (2024)
VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning
por: Xu, Zishan, et al.
Publicado: (2025)
por: Xu, Zishan, et al.
Publicado: (2025)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
por: Villar-Corrales, Angel, et al.
Publicado: (2025)
por: Villar-Corrales, Angel, et al.
Publicado: (2025)
SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs
por: Zhai, Guangyao, et al.
Publicado: (2023)
por: Zhai, Guangyao, et al.
Publicado: (2023)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
por: Zhang, Jiaxu, et al.
Publicado: (2025)
por: Zhang, Jiaxu, et al.
Publicado: (2025)
Automatic Controllable Colorization via Imagination
por: Cong, Xiaoyan, et al.
Publicado: (2024)
por: Cong, Xiaoyan, et al.
Publicado: (2024)
Self-Aware Object Detection via Degradation Manifolds
por: Becker, Stefan, et al.
Publicado: (2026)
por: Becker, Stefan, et al.
Publicado: (2026)
ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis
por: Sun, Zhengwentai, et al.
Publicado: (2026)
por: Sun, Zhengwentai, et al.
Publicado: (2026)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
por: Shaker, Abdelrahman, et al.
Publicado: (2024)
por: Shaker, Abdelrahman, et al.
Publicado: (2024)
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
por: Li, Yajie, et al.
Publicado: (2026)
por: Li, Yajie, et al.
Publicado: (2026)
Object-Centric Latent Action Learning
por: Klepach, Albina, et al.
Publicado: (2025)
por: Klepach, Albina, et al.
Publicado: (2025)
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking
por: Zhou, Xinyu, et al.
Publicado: (2025)
por: Zhou, Xinyu, et al.
Publicado: (2025)
ArtiLatent: Realistic Articulated 3D Object Generation via Structured Latents
por: Chen, Honghua, et al.
Publicado: (2025)
por: Chen, Honghua, et al.
Publicado: (2025)
HDR Video Generation via Latent Alignment with Logarithmic Encoding
por: Korem, Naomi Ken, et al.
Publicado: (2026)
por: Korem, Naomi Ken, et al.
Publicado: (2026)
SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
por: Zhang, Songchun, et al.
Publicado: (2025)
por: Zhang, Songchun, et al.
Publicado: (2025)
Cycle Consistency in Video Object-Centric Learning
por: Zhao, Rongzhen, et al.
Publicado: (2026)
por: Zhao, Rongzhen, et al.
Publicado: (2026)
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
por: Cao, Meng, et al.
Publicado: (2025)
por: Cao, Meng, et al.
Publicado: (2025)
MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World Control
por: Zhou, Enshen, et al.
Publicado: (2024)
por: Zhou, Enshen, et al.
Publicado: (2024)
LTX-Video: Realtime Video Latent Diffusion
por: HaCohen, Yoav, et al.
Publicado: (2024)
por: HaCohen, Yoav, et al.
Publicado: (2024)
Single Image to High-Quality 3D Object via Latent Features
por: Dong, Huanning, et al.
Publicado: (2025)
por: Dong, Huanning, et al.
Publicado: (2025)
VILP: Imitation Learning with Latent Video Planning
por: Xu, Zhengtong, et al.
Publicado: (2025)
por: Xu, Zhengtong, et al.
Publicado: (2025)
Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy
por: Li, You, et al.
Publicado: (2024)
por: Li, You, et al.
Publicado: (2024)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
por: Li, Yian, et al.
Publicado: (2026)
por: Li, Yian, et al.
Publicado: (2026)
Video Generation with Predictive Latents
por: Zhao, Yian, et al.
Publicado: (2026)
por: Zhao, Yian, et al.
Publicado: (2026)
Latent Video Dataset Distillation
por: Li, Ning, et al.
Publicado: (2025)
por: Li, Ning, et al.
Publicado: (2025)
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
por: Chen, Kerui, et al.
Publicado: (2026)
por: Chen, Kerui, et al.
Publicado: (2026)
Mobius: Text to Seamless Looping Video Generation via Latent Shift
por: Bi, Xiuli, et al.
Publicado: (2025)
por: Bi, Xiuli, et al.
Publicado: (2025)
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
por: Wallingford, Matthew, et al.
Publicado: (2024)
por: Wallingford, Matthew, et al.
Publicado: (2024)
LatentColorization: Latent Diffusion-Based Speaker Video Colorization
por: Ward, Rory, et al.
Publicado: (2024)
por: Ward, Rory, et al.
Publicado: (2024)
RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos
por: Xue, Lixin, et al.
Publicado: (2026)
por: Xue, Lixin, et al.
Publicado: (2026)
MOD-UV: Learning Mobile Object Detectors from Unlabeled Videos
por: Sun, Yihong, et al.
Publicado: (2024)
por: Sun, Yihong, et al.
Publicado: (2024)
Ejemplares similares
-
Loci-Segmented: Improving Scene Segmentation Learning
por: Traub, Manuel, et al.
Publicado: (2023) -
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
por: Traub, Manuel, et al.
Publicado: (2025) -
Offline Tracking with Object Permanence
por: Liu, Xianzhong, et al.
Publicado: (2023) -
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
por: Chen, Jiahe, et al.
Publicado: (2026) -
Data-driven Verification of DNNs for Object Recognition
por: Otte, Clemens, et al.
Publicado: (2024)