GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Fuente:
arXiv
Saved in:
| Main Authors: | Hassan, Mariam, Stapf, Sebastian, Rahimi, Ahmad, Rezende, Pedro M B, Haghighi, Yasaman, Brüggemann, David, Katircioglu, Isinsu, Zhang, Lin, Chen, Xiaoran, Saha, Suman, Cannici, Marco, Aljalbout, Elie, Ye, Botao, Wang, Xi, Davtyan, Aram, Salzmann, Mathieu, Scaramuzza, Davide, Pollefeys, Marc, Favaro, Paolo, Alahi, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Communication-Inspired Tokenization for Structured Image Representations
by: Davtyan, Aram, et al.
Published: (2026)
by: Davtyan, Aram, et al.
Published: (2026)
From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models
by: Acuaviva, Pablo, et al.
Published: (2025)
by: Acuaviva, Pablo, et al.
Published: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
by: Acuaviva, Pablo, et al.
Published: (2025)
by: Acuaviva, Pablo, et al.
Published: (2025)
Composition of Memory Experts for Diffusion World Models
by: Stapf, Sebastian, et al.
Published: (2026)
by: Stapf, Sebastian, et al.
Published: (2026)
Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video Generation
by: Davtyan, Aram, et al.
Published: (2023)
by: Davtyan, Aram, et al.
Published: (2023)
Approximate Imitation Learning for Event-based Quadrotor Flight in Cluttered Environments
by: Messikommer, Nico, et al.
Published: (2026)
by: Messikommer, Nico, et al.
Published: (2026)
SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching
by: Haghighi, Yasaman, et al.
Published: (2026)
by: Haghighi, Yasaman, et al.
Published: (2026)
LayerSync: Self-aligning Intermediate Layers
by: Haghighi, Yasaman, et al.
Published: (2025)
by: Haghighi, Yasaman, et al.
Published: (2025)
KOALA++: Efficient Kalman-Based Optimization with Gradient-Covariance Products
by: Xia, Zixuan, et al.
Published: (2025)
by: Xia, Zixuan, et al.
Published: (2025)
Actor-Critic Model Predictive Control: Differentiable Optimization meets Reinforcement Learning for Agile Flight
by: Romero, Angel, et al.
Published: (2023)
by: Romero, Angel, et al.
Published: (2023)
Accelerating Model-Based Reinforcement Learning with State-Space World Models
by: Krinner, Maria, et al.
Published: (2025)
by: Krinner, Maria, et al.
Published: (2025)
Student-Informed Teacher Training
by: Messikommer, Nico, et al.
Published: (2024)
by: Messikommer, Nico, et al.
Published: (2024)
Mitigating Motion Blur in Neural Radiance Fields with Events and Frames
by: Cannici, Marco, et al.
Published: (2024)
by: Cannici, Marco, et al.
Published: (2024)
CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation
by: Davtyan, Aram, et al.
Published: (2024)
by: Davtyan, Aram, et al.
Published: (2024)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
by: Romero, Angel, et al.
Published: (2025)
by: Romero, Angel, et al.
Published: (2025)
Multi-Task Reinforcement Learning for Quadrotors
by: Xing, Jiaxu, et al.
Published: (2024)
by: Xing, Jiaxu, et al.
Published: (2024)
Faster Inference of Flow-Based Generative Models via Improved Data-Noise Coupling
by: Davtyan, Aram, et al.
Published: (2026)
by: Davtyan, Aram, et al.
Published: (2026)
EgoSim: An Egocentric Multi-view Simulator and Real Dataset for Body-worn Cameras during Motion and Activity
by: Hollidt, Dominik, et al.
Published: (2025)
by: Hollidt, Dominik, et al.
Published: (2025)
Ego-Alter Ego
by: Pizer, John
Published: (2020)
by: Pizer, John
Published: (2020)
Event-Aided Sharp Radiance Field Reconstruction for Fast-Flying Drones
by: Zou, Rong, et al.
Published: (2026)
by: Zou, Rong, et al.
Published: (2026)
EgoGen: An Egocentric Synthetic Data Generator
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
EgoM2P: Egocentric Multimodal Multitask Pretraining
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Learning on the Fly: Rapid Policy Adaptation via Differentiable Simulation
by: Pan, Jiahe, et al.
Published: (2025)
by: Pan, Jiahe, et al.
Published: (2025)
EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision
by: Zhao, Yiming, et al.
Published: (2024)
by: Zhao, Yiming, et al.
Published: (2024)
KOALA: A Kalman Optimization Algorithm with Loss Adaptivity
by: Davtyan, Aram, et al.
Published: (2021)
by: Davtyan, Aram, et al.
Published: (2021)
HCQA @ Ego4D EgoSchema Challenge 2024
by: Zhang, Haoyu, et al.
Published: (2024)
by: Zhang, Haoyu, et al.
Published: (2024)
FaVoR: Features via Voxel Rendering for Camera Relocalization
by: Polizzi, Vincenzo, et al.
Published: (2024)
by: Polizzi, Vincenzo, et al.
Published: (2024)
Egocentric Event-Based Vision for Ping Pong Ball Trajectory Prediction
by: Alberico, Ivan, et al.
Published: (2025)
by: Alberico, Ivan, et al.
Published: (2025)
Event-Based De-Snowing for Autonomous Driving
by: Muglikar, Manasi, et al.
Published: (2025)
by: Muglikar, Manasi, et al.
Published: (2025)
EgoExo-WM: Unlocking Exo Video for Ego World Models
by: Tran, Danny, et al.
Published: (2026)
by: Tran, Danny, et al.
Published: (2026)
EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
by: Punamiya, Ryan, et al.
Published: (2025)
by: Punamiya, Ryan, et al.
Published: (2025)
EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos
by: Ma, Junyi, et al.
Published: (2025)
by: Ma, Junyi, et al.
Published: (2025)
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
by: Ge, Mengmeng, et al.
Published: (2026)
by: Ge, Mengmeng, et al.
Published: (2026)
EgoLog: Ego-Centric Fine-Grained Daily Log with Ubiquitous Wearables
by: He, Lixing, et al.
Published: (2025)
by: He, Lixing, et al.
Published: (2025)
HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
by: Wang, Zhi, et al.
Published: (2026)
by: Wang, Zhi, et al.
Published: (2026)
The role of air transportation in emissions through energy usage: Evidence from global data
by: Setareh Katircioglu
Published: (2024)
by: Setareh Katircioglu
Published: (2024)
The effects of an abundance of natural resources on the healthcare industry: The global evidence
by: Setareh Katircioglu
Published: (2024)
by: Setareh Katircioglu
Published: (2024)
A Multi-Loss Strategy for Vehicle Trajectory Prediction: Combining Off-Road, Diversity, and Directional Consistency Losses
by: Rahimi, Ahmad, et al.
Published: (2024)
by: Rahimi, Ahmad, et al.
Published: (2024)
Ego Group Partition: A Novel Framework for Improving Ego Experiments in Social Networks
by: Deng, Lu, et al.
Published: (2024)
by: Deng, Lu, et al.
Published: (2024)
Similar Items
-
Communication-Inspired Tokenization for Structured Image Representations
by: Davtyan, Aram, et al.
Published: (2026) -
From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models
by: Acuaviva, Pablo, et al.
Published: (2025) -
Rethinking Visual Intelligence: Insights from Video Pretraining
by: Acuaviva, Pablo, et al.
Published: (2025) -
Composition of Memory Experts for Diffusion World Models
by: Stapf, Sebastian, et al.
Published: (2026) -
Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video Generation
by: Davtyan, Aram, et al.
Published: (2023)