Walk through Paintings: Egocentric World Models from Internet Priors
Fuente:
arXiv
Saved in:
| Main Authors: | Bagchi, Anurag, Bao, Zhipeng, Bharadhwaj, Homanga, Wang, Yu-Xiong, Tokmakov, Pavel, Hebert, Martial |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReferEverything: Towards Segmenting Everything We Can Speak of in Videos
by: Bagchi, Anurag, et al.
Published: (2024)
by: Bagchi, Anurag, et al.
Published: (2024)
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
by: Bao, Chen, et al.
Published: (2024)
by: Bao, Chen, et al.
Published: (2024)
ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos
by: Soraki, Rustin, et al.
Published: (2026)
by: Soraki, Rustin, et al.
Published: (2026)
Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
by: Zheng, Shuhong, et al.
Published: (2024)
by: Zheng, Shuhong, et al.
Published: (2024)
Separate-and-Enhance: Compositional Finetuning for Text2Image Diffusion Models
by: Bao, Zhipeng, et al.
Published: (2023)
by: Bao, Zhipeng, et al.
Published: (2023)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors
by: Chu, Wen-Hsuan, et al.
Published: (2025)
by: Chu, Wen-Hsuan, et al.
Published: (2025)
Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos
by: Chen, Mingfei, et al.
Published: (2025)
by: Chen, Mingfei, et al.
Published: (2025)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
by: Yun, Heeseung, et al.
Published: (2024)
by: Yun, Heeseung, et al.
Published: (2024)
RoboDream: Compositional World Models for Scalable Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2026)
by: Ye, Junjie, et al.
Published: (2026)
SPIDER: Scalable Physics-Informed Dexterous Retargeting
by: Pan, Chaoyi, et al.
Published: (2025)
by: Pan, Chaoyi, et al.
Published: (2025)
Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction
by: Chen, Hongyi, et al.
Published: (2026)
by: Chen, Hongyi, et al.
Published: (2026)
PlayerOne: Egocentric World Simulator
by: Tu, Yuanpeng, et al.
Published: (2025)
by: Tu, Yuanpeng, et al.
Published: (2025)
Generating Humanless Environment Walkthroughs from Egocentric Walking Tour Videos
by: Ham, Yujin, et al.
Published: (2026)
by: Ham, Yujin, et al.
Published: (2026)
Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
by: Liang, Junbang, et al.
Published: (2024)
by: Liang, Junbang, et al.
Published: (2024)
Semantically Controllable Augmentations for Generalizable Robot Learning
by: Chen, Zoey, et al.
Published: (2024)
by: Chen, Zoey, et al.
Published: (2024)
Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models
by: Chu, Wen-Hsuan, et al.
Published: (2023)
by: Chu, Wen-Hsuan, et al.
Published: (2023)
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
by: Yang, Yue, et al.
Published: (2026)
by: Yang, Yue, et al.
Published: (2026)
Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions
by: Chen, Hongyi, et al.
Published: (2025)
by: Chen, Hongyi, et al.
Published: (2025)
Inverse Painting: Reconstructing The Painting Process
by: Chen, Bowei, et al.
Published: (2024)
by: Chen, Bowei, et al.
Published: (2024)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World
by: Qiu, Heqian, et al.
Published: (2025)
by: Qiu, Heqian, et al.
Published: (2025)
The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text
by: Wang, Hanlin, et al.
Published: (2025)
by: Wang, Hanlin, et al.
Published: (2025)
LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
by: Gao, Quankai, et al.
Published: (2026)
by: Gao, Quankai, et al.
Published: (2026)
EgoAVU: Egocentric Audio-Visual Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
PRevivor: Reviving Ancient Chinese Paintings using Prior-Guided Color Transformers
by: Tang, Tan, et al.
Published: (2025)
by: Tang, Tan, et al.
Published: (2025)
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
Egocentric Whole-Body Human Mesh Recovery with Prior-Guided Learning
by: Na, Soyeon, et al.
Published: (2026)
by: Na, Soyeon, et al.
Published: (2026)
PrefPaint: Enhancing Medical Image Inpainting through Expert Human Feedback
by: Bui, Duy-Bao, et al.
Published: (2025)
by: Bui, Duy-Bao, et al.
Published: (2025)
MADiff: Motion-Aware Mamba Diffusion Models for Hand Trajectory Prediction on Egocentric Videos
by: Ma, Junyi, et al.
Published: (2024)
by: Ma, Junyi, et al.
Published: (2024)
Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis
by: Li, Dayou, et al.
Published: (2026)
by: Li, Dayou, et al.
Published: (2026)
LookOut: Real-World Humanoid Egocentric Navigation
by: Pan, Boxiao, et al.
Published: (2025)
by: Pan, Boxiao, et al.
Published: (2025)
GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusion
by: Guizilini, Vitor, et al.
Published: (2024)
by: Guizilini, Vitor, et al.
Published: (2024)
Computational Approaches for Traditional Chinese Painting: From the "Six Principles of Painting" Perspective
by: Zhang, Wei, et al.
Published: (2023)
by: Zhang, Wei, et al.
Published: (2023)
EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
by: Chen, Lu, et al.
Published: (2025)
by: Chen, Lu, et al.
Published: (2025)
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
PaintScene4D: Consistent 4D Scene Generation from Text Prompts
by: Gupta, Vinayak, et al.
Published: (2024)
by: Gupta, Vinayak, et al.
Published: (2024)
GlobalPaint: Spatiotemporal Coherent Video Outpainting with Global Feature Guidance
by: Pan, Yueming, et al.
Published: (2026)
by: Pan, Yueming, et al.
Published: (2026)
Similar Items
-
ReferEverything: Towards Segmenting Everything We Can Speak of in Videos
by: Bagchi, Anurag, et al.
Published: (2024) -
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024) -
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
by: Bao, Chen, et al.
Published: (2024) -
ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos
by: Soraki, Rustin, et al.
Published: (2026) -
Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
by: Zheng, Shuhong, et al.
Published: (2024)