SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiang, Qiu, Heqian, Wang, Lanxiao, Qiu, Benliu, Meng, Fanman, Xu, Linfeng, Li, Hongliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation
by: Shi, Zhaofeng, et al.
Published: (2025)
by: Shi, Zhaofeng, et al.
Published: (2025)
Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency
by: Shi, Zhaofeng, et al.
Published: (2026)
by: Shi, Zhaofeng, et al.
Published: (2026)
EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World
by: Qiu, Heqian, et al.
Published: (2025)
by: Qiu, Heqian, et al.
Published: (2025)
Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection
by: Gao, Xiangyu, et al.
Published: (2025)
by: Gao, Xiangyu, et al.
Published: (2025)
MCF-VC: Mitigate Catastrophic Forgetting in Class-Incremental Learning for Multimodal Video Captioning
by: Xiong, Huiyu, et al.
Published: (2024)
by: Xiong, Huiyu, et al.
Published: (2024)
Cognition Transferring and Decoupling for Text-supervised Egocentric Semantic Segmentation
by: Shi, Zhaofeng, et al.
Published: (2024)
by: Shi, Zhaofeng, et al.
Published: (2024)
Region Prompt Tuning: Fine-grained Scene Text Detection Utilizing Region Text Prompt
by: Lin, Xingtao, et al.
Published: (2024)
by: Lin, Xingtao, et al.
Published: (2024)
Attention-disentangled Uniform Orthogonal Feature Space Optimization for Few-shot Object Detection
by: Zhao, Taijin, et al.
Published: (2025)
by: Zhao, Taijin, et al.
Published: (2025)
GRSDet: Learning to Generate Local Reverse Samples for Few-shot Object Detection
by: Mei, Hefei, et al.
Published: (2023)
by: Mei, Hefei, et al.
Published: (2023)
Distribution-Level Memory Recall for Continual Learning: Preserving Knowledge and Avoiding Confusion
by: Cheng, Shaoxu, et al.
Published: (2024)
by: Cheng, Shaoxu, et al.
Published: (2024)
Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion
by: Chen, Yukun, et al.
Published: (2025)
by: Chen, Yukun, et al.
Published: (2025)
Challenges and Trends in Egocentric Vision: A Survey
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
by: He, Chiyuan, et al.
Published: (2025)
by: He, Chiyuan, et al.
Published: (2025)
Cross-modal Cognitive Consensus guided Audio-Visual Segmentation
by: Shi, Zhaofeng, et al.
Published: (2023)
by: Shi, Zhaofeng, et al.
Published: (2023)
Continual Learning with Vision-Language Models via Semantic-Geometry Preservation
by: He, Chiyuan, et al.
Published: (2026)
by: He, Chiyuan, et al.
Published: (2026)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
by: Huang, Yifei, et al.
Published: (2024)
by: Huang, Yifei, et al.
Published: (2024)
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
by: Liu, Ruiping, et al.
Published: (2026)
by: Liu, Ruiping, et al.
Published: (2026)
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
by: Jung, Minjoon, et al.
Published: (2025)
by: Jung, Minjoon, et al.
Published: (2025)
FireRescue: A UAV-Based Dataset and Enhanced YOLO Model for Object Detection in Fire Rescue Scenes
by: Xu, Qingyu, et al.
Published: (2025)
by: Xu, Qingyu, et al.
Published: (2025)
MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging
by: Qiu, Zihuan, et al.
Published: (2025)
by: Qiu, Zihuan, et al.
Published: (2025)
Closing the Oracle Gap: Increment Vector Transformation for Class Incremental Learning
by: Qiu, Zihuan, et al.
Published: (2025)
by: Qiu, Zihuan, et al.
Published: (2025)
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
by: Fu, Yuqian, et al.
Published: (2024)
by: Fu, Yuqian, et al.
Published: (2024)
Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction
by: Wei, Dongxu, et al.
Published: (2024)
by: Wei, Dongxu, et al.
Published: (2024)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025)
by: Xu, Jilan, et al.
Published: (2025)
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
by: Reilly, Dominick, et al.
Published: (2025)
by: Reilly, Dominick, et al.
Published: (2025)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
EgoExo-WM: Unlocking Exo Video for Ego World Models
by: Tran, Danny, et al.
Published: (2026)
by: Tran, Danny, et al.
Published: (2026)
Multi-View Industrial Anomaly Detection with Epipolar Constrained Cross-View Fusion
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
ARIC: An Activity Recognition Dataset in Classroom Surveillance Images
by: Xu, Linfeng, et al.
Published: (2024)
by: Xu, Linfeng, et al.
Published: (2024)
MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction
by: Li, Bate, et al.
Published: (2025)
by: Li, Bate, et al.
Published: (2025)
Double Self-weighted Multi-view Clustering via Adaptive View Fusion
by: Fang, Xiang, et al.
Published: (2020)
by: Fang, Xiang, et al.
Published: (2020)
View-Invariant Pixelwise Anomaly Detection in Multi-object Scenes with Adaptive View Synthesis
by: Varghese, Subin, et al.
Published: (2024)
by: Varghese, Subin, et al.
Published: (2024)
Null-Space Filtering for Data-Free Continual Model Merging: Preserving Stability, Promoting Plasticity
by: Qiu, Zihuan, et al.
Published: (2025)
by: Qiu, Zihuan, et al.
Published: (2025)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
by: Gholami, Mohsen, et al.
Published: (2025)
by: Gholami, Mohsen, et al.
Published: (2025)
Match Stereo Videos via Bidirectional Alignment
by: Jing, Junpeng, et al.
Published: (2024)
by: Jing, Junpeng, et al.
Published: (2024)
SAVA: Scalable Learning-Agnostic Data Valuation
by: Kessler, Samuel, et al.
Published: (2024)
by: Kessler, Samuel, et al.
Published: (2024)
EgoTwin: Dreaming Body and View in First Person
by: Xiu, Jingqiao, et al.
Published: (2025)
by: Xiu, Jingqiao, et al.
Published: (2025)
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
by: Li, Yuan-Ming, et al.
Published: (2024)
by: Li, Yuan-Ming, et al.
Published: (2024)
Similar Items
-
Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation
by: Shi, Zhaofeng, et al.
Published: (2025) -
Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency
by: Shi, Zhaofeng, et al.
Published: (2026) -
EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World
by: Qiu, Heqian, et al.
Published: (2025) -
Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection
by: Gao, Xiangyu, et al.
Published: (2025) -
MCF-VC: Mitigate Catastrophic Forgetting in Class-Incremental Learning for Multimodal Video Captioning
by: Xiong, Huiyu, et al.
Published: (2024)