Intention-driven Ego-to-Exo Video Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Luo, Hongchen, Zhu, Kai, Zhai, Wei, Cao, Yang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EgoExo-WM: Unlocking Exo Video for Ego World Models
por: Tran, Danny, et al.
Publicado: (2026)
por: Tran, Danny, et al.
Publicado: (2026)
Bidirectional Progressive Transformer for Interaction Intention Anticipation
por: Zhang, Zichen, et al.
Publicado: (2024)
por: Zhang, Zichen, et al.
Publicado: (2024)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
por: Xu, Jilan, et al.
Publicado: (2025)
por: Xu, Jilan, et al.
Publicado: (2025)
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
por: Shao, Yawen, et al.
Publicado: (2024)
por: Shao, Yawen, et al.
Publicado: (2024)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
por: Huang, Yifei, et al.
Publicado: (2024)
por: Huang, Yifei, et al.
Publicado: (2024)
Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
por: Mahdi, Mohammad, et al.
Publicado: (2025)
por: Mahdi, Mohammad, et al.
Publicado: (2025)
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
por: Mahdi, Mohammad, et al.
Publicado: (2026)
por: Mahdi, Mohammad, et al.
Publicado: (2026)
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
por: Zhang, Haoyu, et al.
Publicado: (2025)
por: Zhang, Haoyu, et al.
Publicado: (2025)
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
por: Liu, Ruiping, et al.
Publicado: (2026)
por: Liu, Ruiping, et al.
Publicado: (2026)
PEAR: Phrase-Based Hand-Object Interaction Anticipation
por: Zhang, Zichen, et al.
Publicado: (2024)
por: Zhang, Zichen, et al.
Publicado: (2024)
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
por: Deng, Huilin, et al.
Publicado: (2024)
por: Deng, Huilin, et al.
Publicado: (2024)
Robust Ego-Exo Correspondence with Long-Term Memory
por: Hu, Yijun, et al.
Publicado: (2025)
por: Hu, Yijun, et al.
Publicado: (2025)
LEMON: Learning 3D Human-Object Interaction Relation from 2D Images
por: Yang, Yuhang, et al.
Publicado: (2023)
por: Yang, Yuhang, et al.
Publicado: (2023)
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
por: Özsoy, Ege, et al.
Publicado: (2025)
por: Özsoy, Ege, et al.
Publicado: (2025)
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
por: Jung, Minjoon, et al.
Publicado: (2025)
por: Jung, Minjoon, et al.
Publicado: (2025)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
por: He, Yuping, et al.
Publicado: (2025)
por: He, Yuping, et al.
Publicado: (2025)
Visual-Geometric Collaborative Guidance for Affordance Learning
por: Luo, Hongchen, et al.
Publicado: (2024)
por: Luo, Hongchen, et al.
Publicado: (2024)
Leverage Task Context for Object Affordance Ranking
por: Huang, Haojie, et al.
Publicado: (2024)
por: Huang, Haojie, et al.
Publicado: (2024)
PCIE_EgoHandPose Solution for EgoExo4D Hand Pose Challenge
por: Chen, Feng, et al.
Publicado: (2024)
por: Chen, Feng, et al.
Publicado: (2024)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
por: Ohkawa, Takehiko, et al.
Publicado: (2023)
por: Ohkawa, Takehiko, et al.
Publicado: (2023)
E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control
por: Gu, Qiao, et al.
Publicado: (2026)
por: Gu, Qiao, et al.
Publicado: (2026)
Visual Context Window Extension: A New Perspective for Long Video Understanding
por: Wei, Hongchen, et al.
Publicado: (2024)
por: Wei, Hongchen, et al.
Publicado: (2024)
ENIGMA-360: An Ego-Exo Dataset for Human Behavior Understanding in Industrial Scenarios
por: Ragusa, Francesco, et al.
Publicado: (2026)
por: Ragusa, Francesco, et al.
Publicado: (2026)
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
por: Li, Yuan-Ming, et al.
Publicado: (2024)
por: Li, Yuan-Ming, et al.
Publicado: (2024)
PCIE_Pose Solution for EgoExo4D Pose and Proficiency Estimation Challenge
por: Chen, Feng, et al.
Publicado: (2025)
por: Chen, Feng, et al.
Publicado: (2025)
CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025
por: Tanoue, Hayato, et al.
Publicado: (2025)
por: Tanoue, Hayato, et al.
Publicado: (2025)
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
por: Ge, Mengmeng, et al.
Publicado: (2026)
por: Ge, Mengmeng, et al.
Publicado: (2026)
ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models
por: Fang, Zixun, et al.
Publicado: (2025)
por: Fang, Zixun, et al.
Publicado: (2025)
SEED4D: A Synthetic Ego--Exo Dynamic 4D Data Generator, Driving Dataset and Benchmark
por: Kästingschäfer, Marius, et al.
Publicado: (2024)
por: Kästingschäfer, Marius, et al.
Publicado: (2024)
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
por: Reilly, Dominick, et al.
Publicado: (2025)
por: Reilly, Dominick, et al.
Publicado: (2025)
Ego-Exo 3D Hand Tracking in the Wild with a Mobile Multi-Camera Rig
por: Rim, Patrick, et al.
Publicado: (2025)
por: Rim, Patrick, et al.
Publicado: (2025)
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
por: Park, Jungin, et al.
Publicado: (2025)
por: Park, Jungin, et al.
Publicado: (2025)
EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views
por: Yang, Yuhang, et al.
Publicado: (2024)
por: Yang, Yuhang, et al.
Publicado: (2024)
EgoLCD: Egocentric Video Generation with Long Context Diffusion
por: Zhang, Liuzhou, et al.
Publicado: (2025)
por: Zhang, Liuzhou, et al.
Publicado: (2025)
HERO: Human Reaction Generation from Videos
por: Yu, Chengjun, et al.
Publicado: (2025)
por: Yu, Chengjun, et al.
Publicado: (2025)
Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
por: Deng, Huilin, et al.
Publicado: (2025)
por: Deng, Huilin, et al.
Publicado: (2025)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
por: Fu, Yuqian, et al.
Publicado: (2025)
por: Fu, Yuqian, et al.
Publicado: (2025)
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
por: Grauman, Kristen, et al.
Publicado: (2023)
por: Grauman, Kristen, et al.
Publicado: (2023)
Grounding 3D Scene Affordance From Egocentric Interactions
por: Liu, Cuiyu, et al.
Publicado: (2024)
por: Liu, Cuiyu, et al.
Publicado: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
por: Wei, Hongchen, et al.
Publicado: (2025)
por: Wei, Hongchen, et al.
Publicado: (2025)
Ejemplares similares
-
EgoExo-WM: Unlocking Exo Video for Ego World Models
por: Tran, Danny, et al.
Publicado: (2026) -
Bidirectional Progressive Transformer for Interaction Intention Anticipation
por: Zhang, Zichen, et al.
Publicado: (2024) -
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
por: Xu, Jilan, et al.
Publicado: (2025) -
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
por: Shao, Yawen, et al.
Publicado: (2024) -
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
por: Huang, Yifei, et al.
Publicado: (2024)