Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
Fuente:
arXiv
Guardado en:
| Autores principales: | Gu, Yuming, Wang, Yizhi, Hong, Yining, Gao, Yipeng, Jiang, Hao, Wang, Angtian, Liu, Bo, Dennler, Nathaniel S., Kuang, Zhengfei, Li, Hao, Wetzstein, Gordon, Ma, Chongyang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
por: Kuang, Zhengfei, et al.
Publicado: (2024)
por: Kuang, Zhengfei, et al.
Publicado: (2024)
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
por: Ackermann, Jan, et al.
Publicado: (2026)
por: Ackermann, Jan, et al.
Publicado: (2026)
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
por: Cong, Xiaoyan, et al.
Publicado: (2025)
por: Cong, Xiaoyan, et al.
Publicado: (2025)
VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
por: Kuang, Zhengfei, et al.
Publicado: (2025)
por: Kuang, Zhengfei, et al.
Publicado: (2025)
Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors
por: Kuang, Zhengfei, et al.
Publicado: (2024)
por: Kuang, Zhengfei, et al.
Publicado: (2024)
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
por: Deng, Yufan, et al.
Publicado: (2025)
por: Deng, Yufan, et al.
Publicado: (2025)
ATI: Any Trajectory Instruction for Controllable Video Generation
por: Wang, Angtian, et al.
Publicado: (2025)
por: Wang, Angtian, et al.
Publicado: (2025)
BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
por: Wang, Yiming, et al.
Publicado: (2025)
por: Wang, Yiming, et al.
Publicado: (2025)
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
por: Deng, Yufan, et al.
Publicado: (2025)
por: Deng, Yufan, et al.
Publicado: (2025)
Spectral Progressive Diffusion for Efficient Image and Video Generation
por: Xiao, Howard, et al.
Publicado: (2026)
por: Xiao, Howard, et al.
Publicado: (2026)
Foveated Diffusion: Efficient Spatially Adaptive Image and Video Generation
por: Chao, Brian, et al.
Publicado: (2026)
por: Chao, Brian, et al.
Publicado: (2026)
BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
por: Po, Ryan, et al.
Publicado: (2025)
por: Po, Ryan, et al.
Publicado: (2025)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
por: He, Hao, et al.
Publicado: (2024)
por: He, Hao, et al.
Publicado: (2024)
HECTOR: Hybrid Editable Compositional Object References for Video Generation
por: Zhang, Guofeng, et al.
Publicado: (2026)
por: Zhang, Guofeng, et al.
Publicado: (2026)
Contrastive Learning from Exploratory Actions: Leveraging Natural Interactions for Preference Elicitation
por: Dennler, Nathaniel, et al.
Publicado: (2025)
por: Dennler, Nathaniel, et al.
Publicado: (2025)
GIFT: Generalizing Intent for Flexible Test-Time Rewards
por: Amin, Fin, et al.
Publicado: (2026)
por: Amin, Fin, et al.
Publicado: (2026)
Using Causal Trees to Estimate Personalized Task Difficulty in Post-Stroke Individuals
por: Dennler, Nathaniel, et al.
Publicado: (2024)
por: Dennler, Nathaniel, et al.
Publicado: (2024)
The Current State of AI Bias Bounties: An Overview of Existing Programmes and Research
por: Kucenko, Sergej, et al.
Publicado: (2025)
por: Kucenko, Sergej, et al.
Publicado: (2025)
Singing the Body Electric: The Impact of Robot Embodiment on User Expectations
por: Dennler, Nathaniel, et al.
Publicado: (2024)
por: Dennler, Nathaniel, et al.
Publicado: (2024)
Dual Ascent Diffusion for Inverse Problems
por: Kim, Minseo, et al.
Publicado: (2025)
por: Kim, Minseo, et al.
Publicado: (2025)
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
por: Zhang, Lvmin, et al.
Publicado: (2025)
por: Zhang, Lvmin, et al.
Publicado: (2025)
Generic 3D Diffusion Adapter Using Controlled Multi-View Editing
por: Chen, Hansheng, et al.
Publicado: (2024)
por: Chen, Hansheng, et al.
Publicado: (2024)
Turning Text and Imagery into Captivating Visual Video
por: Wang, Mingming, et al.
Publicado: (2024)
por: Wang, Mingming, et al.
Publicado: (2024)
Infinite Gaze Generation for Videos with Autoregressive Diffusion
por: Kang, Jenna, et al.
Publicado: (2026)
por: Kang, Jenna, et al.
Publicado: (2026)
Position: Olfaction Standardization is Essential for the Advancement of Embodied Artificial Intelligence
por: France, Kordel K., et al.
Publicado: (2025)
por: France, Kordel K., et al.
Publicado: (2025)
Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
por: Deng, Boyang, et al.
Publicado: (2024)
por: Deng, Boyang, et al.
Publicado: (2024)
EmbodiedPlace: Learning Mixture-of-Features with Embodied Constraints for Visual Place Recognition
por: Liu, Bingxi, et al.
Publicado: (2025)
por: Liu, Bingxi, et al.
Publicado: (2025)
Orthogonal Adaptation for Modular Customization of Diffusion Models
por: Po, Ryan, et al.
Publicado: (2023)
por: Po, Ryan, et al.
Publicado: (2023)
X-Dyna: Expressive Dynamic Human Image Animation
por: Chang, Di, et al.
Publicado: (2025)
por: Chang, Di, et al.
Publicado: (2025)
Designing Robot Identity: The Role of Voice, Clothing, and Task on Robot Gender Perception
por: Dennler, Nathaniel S., et al.
Publicado: (2024)
por: Dennler, Nathaniel S., et al.
Publicado: (2024)
Improving User Experience in Preference-Based Optimization of Reward Functions for Assistive Robots
por: Dennler, Nathaniel, et al.
Publicado: (2024)
por: Dennler, Nathaniel, et al.
Publicado: (2024)
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
por: Zhang, Guofeng, et al.
Publicado: (2025)
por: Zhang, Guofeng, et al.
Publicado: (2025)
EnerVerse-AC: Envisioning Embodied Environments with Action Condition
por: Jiang, Yuxin, et al.
Publicado: (2025)
por: Jiang, Yuxin, et al.
Publicado: (2025)
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
por: Huang, Siyuan, et al.
Publicado: (2025)
por: Huang, Siyuan, et al.
Publicado: (2025)
LiPUP-MA: A Residential Experience-centric Multi-Agent Framework for Living-in-the-loop Participatory Urban Planning
por: Ni, Hang, et al.
Publicado: (2024)
por: Ni, Hang, et al.
Publicado: (2024)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
por: Fang, Ye, et al.
Publicado: (2025)
por: Fang, Ye, et al.
Publicado: (2025)
CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models
por: He, Hao, et al.
Publicado: (2025)
por: He, Hao, et al.
Publicado: (2025)
Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL
por: Zhong, Fangwei, et al.
Publicado: (2024)
por: Zhong, Fangwei, et al.
Publicado: (2024)
Hierarchical Instruction-aware Embodied Visual Tracking
por: Wu, Kui, et al.
Publicado: (2025)
por: Wu, Kui, et al.
Publicado: (2025)
Neural Ganglion Sensors: Learning Task-specific Event Cameras Inspired by the Neural Circuit of the Human Retina
por: So, Haley M., et al.
Publicado: (2025)
por: So, Haley M., et al.
Publicado: (2025)
Ejemplares similares
-
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
por: Kuang, Zhengfei, et al.
Publicado: (2024) -
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
por: Ackermann, Jan, et al.
Publicado: (2026) -
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
por: Cong, Xiaoyan, et al.
Publicado: (2025) -
VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
por: Kuang, Zhengfei, et al.
Publicado: (2025) -
Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors
por: Kuang, Zhengfei, et al.
Publicado: (2024)