TextOCVP: Object-Centric Video Prediction with Language Guidance
Fuente:
arXiv
Guardado en:
| Autores principales: | Villar-Corrales, Angel, Plepi, Gjergj, Behnke, Sven |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
por: Villar-Corrales, Angel, et al.
Publicado: (2025)
por: Villar-Corrales, Angel, et al.
Publicado: (2025)
MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion
por: Villar-Corrales, Angel, et al.
Publicado: (2024)
por: Villar-Corrales, Angel, et al.
Publicado: (2024)
VideoPCDNet: Video Parsing and Prediction with Phase Correlation Networks
por: Vicente, Noel José Rodrigues, et al.
Publicado: (2025)
por: Vicente, Noel José Rodrigues, et al.
Publicado: (2025)
OC-SOP: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness
por: Cao, Helin, et al.
Publicado: (2025)
por: Cao, Helin, et al.
Publicado: (2025)
Top-Down Guidance for Learning Object-Centric Representations
por: Zou, Junhong, et al.
Publicado: (2024)
por: Zou, Junhong, et al.
Publicado: (2024)
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
por: Pätzold, Bastian, et al.
Publicado: (2025)
por: Pätzold, Bastian, et al.
Publicado: (2025)
Learning Embeddings with Centroid Triplet Loss for Object Identification in Robotic Grasping
por: Gouda, Anas, et al.
Publicado: (2024)
por: Gouda, Anas, et al.
Publicado: (2024)
TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
por: Wang, Xingrui, et al.
Publicado: (2024)
por: Wang, Xingrui, et al.
Publicado: (2024)
HyenaPixel: Global Image Context with Convolutions
por: Spravil, Julian, et al.
Publicado: (2024)
por: Spravil, Julian, et al.
Publicado: (2024)
FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-pose, and Facial Expression Features
por: Rochow, Andre, et al.
Publicado: (2024)
por: Rochow, Andre, et al.
Publicado: (2024)
Object-Centric Framework for Video Moment Retrieval
por: Li, Zongyao, et al.
Publicado: (2025)
por: Li, Zongyao, et al.
Publicado: (2025)
Cycle Consistency in Video Object-Centric Learning
por: Zhao, Rongzhen, et al.
Publicado: (2026)
por: Zhao, Rongzhen, et al.
Publicado: (2026)
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
por: Zhou, Jiaying, et al.
Publicado: (2026)
por: Zhou, Jiaying, et al.
Publicado: (2026)
SWA-SOP: Spatially-aware Window Attention for Semantic Occupancy Prediction in Autonomous Driving
por: Cao, Helin, et al.
Publicado: (2025)
por: Cao, Helin, et al.
Publicado: (2025)
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
por: Song, Yeon-Ji, et al.
Publicado: (2024)
por: Song, Yeon-Ji, et al.
Publicado: (2024)
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
por: Wang, Yihao, et al.
Publicado: (2025)
por: Wang, Yihao, et al.
Publicado: (2025)
Compositional Video Synthesis by Temporal Object-Centric Learning
por: Akan, Adil Kaan, et al.
Publicado: (2025)
por: Akan, Adil Kaan, et al.
Publicado: (2025)
DiffSSC: Semantic LiDAR Scan Completion using Denoising Diffusion Probabilistic Models
por: Cao, Helin, et al.
Publicado: (2024)
por: Cao, Helin, et al.
Publicado: (2024)
SLCF-Net: Sequential LiDAR-Camera Fusion for Semantic Scene Completion using a 3D Recurrent U-Net
por: Cao, Helin, et al.
Publicado: (2024)
por: Cao, Helin, et al.
Publicado: (2024)
Learning from SAM: Harnessing a Foundation Model for Sim2Real Adaptation by Regularization
por: Bonani, Mayara E., et al.
Publicado: (2023)
por: Bonani, Mayara E., et al.
Publicado: (2023)
Marker-free Human Gait Analysis using a Smart Edge Sensor System
por: Bauer, Eva Katharina, et al.
Publicado: (2024)
por: Bauer, Eva Katharina, et al.
Publicado: (2024)
Object-Aware Video Matting with Cross-Frame Guidance
por: Zhang, Huayu, et al.
Publicado: (2025)
por: Zhang, Huayu, et al.
Publicado: (2025)
Feature-Preserving Mesh Decimation for Normal Integration
por: Heep, Moritz, et al.
Publicado: (2025)
por: Heep, Moritz, et al.
Publicado: (2025)
GenEraser: Generalizable Video Object Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver
por: Chen, Yuqing, et al.
Publicado: (2026)
por: Chen, Yuqing, et al.
Publicado: (2026)
Video Spatial Reasoning with Object-Centric 3D Rollout
por: Tang, Haoran, et al.
Publicado: (2025)
por: Tang, Haoran, et al.
Publicado: (2025)
VASE: Object-Centric Appearance and Shape Manipulation of Real Videos
por: Peruzzo, Elia, et al.
Publicado: (2024)
por: Peruzzo, Elia, et al.
Publicado: (2024)
Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations
por: Spieler, Jonathan, et al.
Publicado: (2026)
por: Spieler, Jonathan, et al.
Publicado: (2026)
ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
por: Guo, Wenliang, et al.
Publicado: (2025)
por: Guo, Wenliang, et al.
Publicado: (2025)
HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
por: Li, Yayuan, et al.
Publicado: (2024)
por: Li, Yayuan, et al.
Publicado: (2024)
DDLP: Unsupervised Object-Centric Video Prediction with Deep Dynamic Latent Particles
por: Daniel, Tal, et al.
Publicado: (2023)
por: Daniel, Tal, et al.
Publicado: (2023)
Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
por: Zadaianchuk, Andrii, et al.
Publicado: (2023)
por: Zadaianchuk, Andrii, et al.
Publicado: (2023)
Iterative Motion Compensation for Canonical 3D Reconstruction from UAV Plant Images Captured in Windy Conditions
por: Rochow, Andre, et al.
Publicado: (2025)
por: Rochow, Andre, et al.
Publicado: (2025)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
por: Han, Jiwook, et al.
Publicado: (2026)
por: Han, Jiwook, et al.
Publicado: (2026)
Object-Centric Diffusion for Efficient Video Editing
por: Kahatapitiya, Kumara, et al.
Publicado: (2024)
por: Kahatapitiya, Kumara, et al.
Publicado: (2024)
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
por: Oh, Yoonjin, et al.
Publicado: (2025)
por: Oh, Yoonjin, et al.
Publicado: (2025)
TRACE: Object Motion Editing in Videos with First-Frame Trajectory Guidance
por: Phung, Quynh, et al.
Publicado: (2026)
por: Phung, Quynh, et al.
Publicado: (2026)
Efficient Image Annotation via Semi-Supervised Object Segmentation with Label Propagation
por: Tutevych, Vitalii, et al.
Publicado: (2026)
por: Tutevych, Vitalii, et al.
Publicado: (2026)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
por: Wang, Cong, et al.
Publicado: (2023)
por: Wang, Cong, et al.
Publicado: (2023)
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
por: Ge, Mengmeng, et al.
Publicado: (2026)
por: Ge, Mengmeng, et al.
Publicado: (2026)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
por: Dou, Weijia, et al.
Publicado: (2026)
por: Dou, Weijia, et al.
Publicado: (2026)
Ejemplares similares
-
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
por: Villar-Corrales, Angel, et al.
Publicado: (2025) -
MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion
por: Villar-Corrales, Angel, et al.
Publicado: (2024) -
VideoPCDNet: Video Parsing and Prediction with Phase Correlation Networks
por: Vicente, Noel José Rodrigues, et al.
Publicado: (2025) -
OC-SOP: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness
por: Cao, Helin, et al.
Publicado: (2025) -
Top-Down Guidance for Learning Object-Centric Representations
por: Zou, Junhong, et al.
Publicado: (2024)