Learning Object Permanence from Videos via Latent Imaginations
Fuente:
arXiv
Salvato in:
| Autori principali: | Traub, Manuel, Becker, Frederic, Otte, Sebastian, Butz, Martin V. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Loci-Segmented: Improving Scene Segmentation Learning
di: Traub, Manuel, et al.
Pubblicazione: (2023)
di: Traub, Manuel, et al.
Pubblicazione: (2023)
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
di: Traub, Manuel, et al.
Pubblicazione: (2025)
di: Traub, Manuel, et al.
Pubblicazione: (2025)
Offline Tracking with Object Permanence
di: Liu, Xianzhong, et al.
Pubblicazione: (2023)
di: Liu, Xianzhong, et al.
Pubblicazione: (2023)
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
di: Chen, Jiahe, et al.
Pubblicazione: (2026)
di: Chen, Jiahe, et al.
Pubblicazione: (2026)
Data-driven Verification of DNNs for Object Recognition
di: Otte, Clemens, et al.
Pubblicazione: (2024)
di: Otte, Clemens, et al.
Pubblicazione: (2024)
Detection of Fast-Moving Objects with Neuromorphic Hardware
di: Ziegler, Andreas, et al.
Pubblicazione: (2024)
di: Ziegler, Andreas, et al.
Pubblicazione: (2024)
Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation
di: Lou, Zijie, et al.
Pubblicazione: (2026)
di: Lou, Zijie, et al.
Pubblicazione: (2026)
Imagining the Unseen: Generative Location Modeling for Object Placement
di: Yun, Jooyeol, et al.
Pubblicazione: (2024)
di: Yun, Jooyeol, et al.
Pubblicazione: (2024)
Jigsaw++: Imagining Complete Shape Priors for Object Reassembly
di: Lu, Jiaxin, et al.
Pubblicazione: (2024)
di: Lu, Jiaxin, et al.
Pubblicazione: (2024)
Imagine360: Immersive 360 Video Generation from Perspective Anchor
di: Tan, Jing, et al.
Pubblicazione: (2024)
di: Tan, Jing, et al.
Pubblicazione: (2024)
VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning
di: Xu, Zishan, et al.
Pubblicazione: (2025)
di: Xu, Zishan, et al.
Pubblicazione: (2025)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)
SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs
di: Zhai, Guangyao, et al.
Pubblicazione: (2023)
di: Zhai, Guangyao, et al.
Pubblicazione: (2023)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
di: Zhang, Jiaxu, et al.
Pubblicazione: (2025)
di: Zhang, Jiaxu, et al.
Pubblicazione: (2025)
Automatic Controllable Colorization via Imagination
di: Cong, Xiaoyan, et al.
Pubblicazione: (2024)
di: Cong, Xiaoyan, et al.
Pubblicazione: (2024)
Self-Aware Object Detection via Degradation Manifolds
di: Becker, Stefan, et al.
Pubblicazione: (2026)
di: Becker, Stefan, et al.
Pubblicazione: (2026)
ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis
di: Sun, Zhengwentai, et al.
Pubblicazione: (2026)
di: Sun, Zhengwentai, et al.
Pubblicazione: (2026)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2024)
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2024)
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
di: Li, Yajie, et al.
Pubblicazione: (2026)
di: Li, Yajie, et al.
Pubblicazione: (2026)
Object-Centric Latent Action Learning
di: Klepach, Albina, et al.
Pubblicazione: (2025)
di: Klepach, Albina, et al.
Pubblicazione: (2025)
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking
di: Zhou, Xinyu, et al.
Pubblicazione: (2025)
di: Zhou, Xinyu, et al.
Pubblicazione: (2025)
ArtiLatent: Realistic Articulated 3D Object Generation via Structured Latents
di: Chen, Honghua, et al.
Pubblicazione: (2025)
di: Chen, Honghua, et al.
Pubblicazione: (2025)
HDR Video Generation via Latent Alignment with Logarithmic Encoding
di: Korem, Naomi Ken, et al.
Pubblicazione: (2026)
di: Korem, Naomi Ken, et al.
Pubblicazione: (2026)
SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
di: Zhang, Songchun, et al.
Pubblicazione: (2025)
di: Zhang, Songchun, et al.
Pubblicazione: (2025)
Cycle Consistency in Video Object-Centric Learning
di: Zhao, Rongzhen, et al.
Pubblicazione: (2026)
di: Zhao, Rongzhen, et al.
Pubblicazione: (2026)
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
di: Cao, Meng, et al.
Pubblicazione: (2025)
di: Cao, Meng, et al.
Pubblicazione: (2025)
MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World Control
di: Zhou, Enshen, et al.
Pubblicazione: (2024)
di: Zhou, Enshen, et al.
Pubblicazione: (2024)
LTX-Video: Realtime Video Latent Diffusion
di: HaCohen, Yoav, et al.
Pubblicazione: (2024)
di: HaCohen, Yoav, et al.
Pubblicazione: (2024)
Single Image to High-Quality 3D Object via Latent Features
di: Dong, Huanning, et al.
Pubblicazione: (2025)
di: Dong, Huanning, et al.
Pubblicazione: (2025)
VILP: Imitation Learning with Latent Video Planning
di: Xu, Zhengtong, et al.
Pubblicazione: (2025)
di: Xu, Zhengtong, et al.
Pubblicazione: (2025)
Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy
di: Li, You, et al.
Pubblicazione: (2024)
di: Li, You, et al.
Pubblicazione: (2024)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
di: Li, Yian, et al.
Pubblicazione: (2026)
di: Li, Yian, et al.
Pubblicazione: (2026)
Video Generation with Predictive Latents
di: Zhao, Yian, et al.
Pubblicazione: (2026)
di: Zhao, Yian, et al.
Pubblicazione: (2026)
Latent Video Dataset Distillation
di: Li, Ning, et al.
Pubblicazione: (2025)
di: Li, Ning, et al.
Pubblicazione: (2025)
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
di: Chen, Kerui, et al.
Pubblicazione: (2026)
di: Chen, Kerui, et al.
Pubblicazione: (2026)
Mobius: Text to Seamless Looping Video Generation via Latent Shift
di: Bi, Xiuli, et al.
Pubblicazione: (2025)
di: Bi, Xiuli, et al.
Pubblicazione: (2025)
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
di: Wallingford, Matthew, et al.
Pubblicazione: (2024)
di: Wallingford, Matthew, et al.
Pubblicazione: (2024)
LatentColorization: Latent Diffusion-Based Speaker Video Colorization
di: Ward, Rory, et al.
Pubblicazione: (2024)
di: Ward, Rory, et al.
Pubblicazione: (2024)
RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos
di: Xue, Lixin, et al.
Pubblicazione: (2026)
di: Xue, Lixin, et al.
Pubblicazione: (2026)
MOD-UV: Learning Mobile Object Detectors from Unlabeled Videos
di: Sun, Yihong, et al.
Pubblicazione: (2024)
di: Sun, Yihong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Loci-Segmented: Improving Scene Segmentation Learning
di: Traub, Manuel, et al.
Pubblicazione: (2023) -
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
di: Traub, Manuel, et al.
Pubblicazione: (2025) -
Offline Tracking with Object Permanence
di: Liu, Xianzhong, et al.
Pubblicazione: (2023) -
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
di: Chen, Jiahe, et al.
Pubblicazione: (2026) -
Data-driven Verification of DNNs for Object Recognition
di: Otte, Clemens, et al.
Pubblicazione: (2024)