From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
Fuente:
arXiv
Saved in:
| Main Authors: | Reilly, Dominick, Govind, Manish Kumar, Xue, Le, Das, Srijan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
by: Reilly, Dominick, et al.
Published: (2025)
by: Reilly, Dominick, et al.
Published: (2025)
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
by: Reilly, Dominick, et al.
Published: (2024)
by: Reilly, Dominick, et al.
Published: (2024)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
by: Govind, Manish Kumar, et al.
Published: (2026)
by: Govind, Manish Kumar, et al.
Published: (2026)
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads
by: Rahimian, Ali K., et al.
Published: (2024)
by: Rahimian, Ali K., et al.
Published: (2024)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
by: Huang, Yifei, et al.
Published: (2024)
by: Huang, Yifei, et al.
Published: (2024)
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
by: Özsoy, Ege, et al.
Published: (2025)
by: Özsoy, Ege, et al.
Published: (2025)
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
by: Jung, Minjoon, et al.
Published: (2025)
by: Jung, Minjoon, et al.
Published: (2025)
EgoExo-WM: Unlocking Exo Video for Ego World Models
by: Tran, Danny, et al.
Published: (2026)
by: Tran, Danny, et al.
Published: (2026)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025)
by: Xu, Jilan, et al.
Published: (2025)
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
by: Mahdi, Mohammad, et al.
Published: (2026)
by: Mahdi, Mohammad, et al.
Published: (2026)
Gaze-Regularized VLMs for Ego-Centric Behavior Understanding
by: Pani, Anupam, et al.
Published: (2026)
by: Pani, Anupam, et al.
Published: (2026)
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
by: Grauman, Kristen, et al.
Published: (2023)
by: Grauman, Kristen, et al.
Published: (2023)
ENIGMA-360: An Ego-Exo Dataset for Human Behavior Understanding in Industrial Scenarios
by: Ragusa, Francesco, et al.
Published: (2026)
by: Ragusa, Francesco, et al.
Published: (2026)
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
Intention-driven Ego-to-Exo Video Generation
by: Luo, Hongchen, et al.
Published: (2024)
by: Luo, Hongchen, et al.
Published: (2024)
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
by: Liu, Ruiping, et al.
Published: (2026)
by: Liu, Ruiping, et al.
Published: (2026)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
by: Fu, Yuqian, et al.
Published: (2024)
by: Fu, Yuqian, et al.
Published: (2024)
Robust Ego-Exo Correspondence with Long-Term Memory
by: Hu, Yijun, et al.
Published: (2025)
by: Hu, Yijun, et al.
Published: (2025)
SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
by: Li, Yuan-Ming, et al.
Published: (2024)
by: Li, Yuan-Ming, et al.
Published: (2024)
PCIE_EgoHandPose Solution for EgoExo4D Hand Pose Challenge
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
by: Ohkawa, Takehiko, et al.
Published: (2023)
by: Ohkawa, Takehiko, et al.
Published: (2023)
PCIE_Pose Solution for EgoExo4D Pose and Proficiency Estimation Challenge
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025
by: Tanoue, Hayato, et al.
Published: (2025)
by: Tanoue, Hayato, et al.
Published: (2025)
Abductive Ego-View Accident Video Understanding for Safe Driving Perception
by: Fang, Jianwu, et al.
Published: (2024)
by: Fang, Jianwu, et al.
Published: (2024)
Ego-Exo 3D Hand Tracking in the Wild with a Mobile Multi-Camera Rig
by: Rim, Patrick, et al.
Published: (2025)
by: Rim, Patrick, et al.
Published: (2025)
MyVLM: Personalizing VLMs for User-Specific Queries
by: Alaluf, Yuval, et al.
Published: (2024)
by: Alaluf, Yuval, et al.
Published: (2024)
Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
by: Mahdi, Mohammad, et al.
Published: (2025)
by: Mahdi, Mohammad, et al.
Published: (2025)
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier
by: Howlader, Prantik, et al.
Published: (2024)
by: Howlader, Prantik, et al.
Published: (2024)
Fusion-SSAT: Unleashing the Potential of Self-supervised Auxiliary Task by Feature Fusion for Generalized Deepfake Detection
by: Reddy, Shukesh, et al.
Published: (2026)
by: Reddy, Shukesh, et al.
Published: (2026)
EgoAVU: Egocentric Audio-Visual Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
SEED4D: A Synthetic Ego--Exo Dynamic 4D Data Generator, Driving Dataset and Benchmark
by: Kästingschäfer, Marius, et al.
Published: (2024)
by: Kästingschäfer, Marius, et al.
Published: (2024)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
by: Zhu, Bingwen, et al.
Published: (2026)
by: Zhu, Bingwen, et al.
Published: (2026)
Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency
by: Shi, Zhaofeng, et al.
Published: (2026)
by: Shi, Zhaofeng, et al.
Published: (2026)
Introducing Gating and Context into Temporal Action Detection
by: Reka, Aglind, et al.
Published: (2024)
by: Reka, Aglind, et al.
Published: (2024)
Similar Items
-
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
by: Reilly, Dominick, et al.
Published: (2025) -
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
by: Reilly, Dominick, et al.
Published: (2024) -
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
by: Govind, Manish Kumar, et al.
Published: (2026) -
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
by: Sinha, Arkaprava, et al.
Published: (2025) -
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads
by: Rahimian, Ali K., et al.
Published: (2024)