Whole-Body Conditioned Egocentric Video Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Yutong, Tran, Danny, Bar, Amir, LeCun, Yann, Darrell, Trevor, Malik, Jitendra |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
EgoPet: Egomotion and Interaction Data from an Animal's Perspective
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
Hierarchical World Models as Visual Whole-Body Humanoid Controllers
by: Hansen, Nicklas, et al.
Published: (2024)
by: Hansen, Nicklas, et al.
Published: (2024)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
EVA: An Embodied World Model for Future Video Anticipation
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos
by: Zhu, Xilei, et al.
Published: (2024)
by: Zhu, Xilei, et al.
Published: (2024)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
HR-INR: Continuous Space-Time Video Super-Resolution via Event Camera
by: Lu, Yunfan, et al.
Published: (2024)
by: Lu, Yunfan, et al.
Published: (2024)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
WoW: Towards a World omniscient World model Through Embodied Interaction
by: Chi, Xiaowei, et al.
Published: (2025)
by: Chi, Xiaowei, et al.
Published: (2025)
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
by: Zheng, Jiayi, et al.
Published: (2025)
by: Zheng, Jiayi, et al.
Published: (2025)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
by: Min, Chen, et al.
Published: (2023)
by: Min, Chen, et al.
Published: (2023)
XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments
by: Qian, Kangan, et al.
Published: (2026)
by: Qian, Kangan, et al.
Published: (2026)
Integrating Multi-Modal Sensors: A Review of Fusion Techniques for Intelligent Vehicles
by: Wei, Chuheng, et al.
Published: (2025)
by: Wei, Chuheng, et al.
Published: (2025)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
by: Wang, Sen, et al.
Published: (2025)
by: Wang, Sen, et al.
Published: (2025)
DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
by: Zou, Jiayi, et al.
Published: (2025)
by: Zou, Jiayi, et al.
Published: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video Synthesis
by: Fang, Jianwu, et al.
Published: (2025)
by: Fang, Jianwu, et al.
Published: (2025)
ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial Images
by: Zhu, Xilei, et al.
Published: (2024)
by: Zhu, Xilei, et al.
Published: (2024)
Finding Visual Task Vectors
by: Hojel, Alberto, et al.
Published: (2024)
by: Hojel, Alberto, et al.
Published: (2024)
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
by: Ma, Jianzhe, et al.
Published: (2026)
by: Ma, Jianzhe, et al.
Published: (2026)
RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
by: Peng, Kunyu, et al.
Published: (2025)
by: Peng, Kunyu, et al.
Published: (2025)
A Very Big Video Reasoning Suite
by: Wang, Maijunxian, et al.
Published: (2026)
by: Wang, Maijunxian, et al.
Published: (2026)
CIV-DG: Conditional Instrumental Variables for Domain Generalization in Medical Imaging
by: Bai, Shaojin, et al.
Published: (2026)
by: Bai, Shaojin, et al.
Published: (2026)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
by: Li, Jie, et al.
Published: (2023)
by: Li, Jie, et al.
Published: (2023)
StereoVAE: A lightweight stereo-matching system using embedded GPUs
by: Chang, Qiong, et al.
Published: (2023)
by: Chang, Qiong, et al.
Published: (2023)
Customizable Perturbation Synthesis for Robust SLAM Benchmarking
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
Stochastic positional embeddings improve masked image modeling
by: Bar, Amir, et al.
Published: (2023)
by: Bar, Amir, et al.
Published: (2023)
VideoMem: Constructing, Analyzing, Predicting Short-term and Long-term Video Memorability
by: Cohendet, Romain, et al.
Published: (2018)
by: Cohendet, Romain, et al.
Published: (2018)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
Advanced Learning-Based Inter Prediction for Future Video Coding
by: Zhao, Yanchen, et al.
Published: (2024)
by: Zhao, Yanchen, et al.
Published: (2024)
NavigScene: Bridging Local Perception and Global Navigation for Beyond-Visual-Range Autonomous Driving
by: Peng, Qucheng, et al.
Published: (2025)
by: Peng, Qucheng, et al.
Published: (2025)
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
by: Moradi, Morteza, et al.
Published: (2024)
by: Moradi, Morteza, et al.
Published: (2024)
Similar Items
-
Navigation World Models
by: Bar, Amir, et al.
Published: (2024) -
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025) -
EgoPet: Egomotion and Interaction Data from an Animal's Perspective
by: Bar, Amir, et al.
Published: (2024) -
Hierarchical World Models as Visual Whole-Body Humanoid Controllers
by: Hansen, Nicklas, et al.
Published: (2024) -
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)