ViewBridge: Curriculum Knowledge Distillation for Activity View-Invariance Under Extreme Viewpoint Changes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Somayazulu, Arjun, Mavroudi, Efi, Chen, Changan, Torresani, Lorenzo, Grauman, Kristen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2026)
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2026)
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2024)
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2024)
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
Few-View Object Reconstruction with Unknown Categories and Camera Poses
von: Jiang, Hanwen, et al.
Veröffentlicht: (2022)
von: Jiang, Hanwen, et al.
Veröffentlicht: (2022)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
von: An, Joungbin, et al.
Veröffentlicht: (2025)
von: An, Joungbin, et al.
Veröffentlicht: (2025)
Learning Object State Changes in Videos: An Open-World Perspective
von: Xue, Zihui, et al.
Veröffentlicht: (2023)
von: Xue, Zihui, et al.
Veröffentlicht: (2023)
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
von: Chen, Changan, et al.
Veröffentlicht: (2024)
von: Chen, Changan, et al.
Veröffentlicht: (2024)
Learning Skill-Attributes for Transferable Assessment in Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025)
Cross-View Consistency Regularisation for Knowledge Distillation
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
Step Differences in Instructional Video
von: Nagarajan, Tushar, et al.
Veröffentlicht: (2024)
von: Nagarajan, Tushar, et al.
Veröffentlicht: (2024)
Triple-View Knowledge Distillation for Semi-Supervised Semantic Segmentation
von: Li, Ping, et al.
Veröffentlicht: (2023)
von: Li, Ping, et al.
Veröffentlicht: (2023)
BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation
von: Kim, Beomjun, et al.
Veröffentlicht: (2025)
von: Kim, Beomjun, et al.
Veröffentlicht: (2025)
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
von: An, Joungbin, et al.
Veröffentlicht: (2026)
von: An, Joungbin, et al.
Veröffentlicht: (2026)
SPOC: Spatially-Progressing Object State Change Segmentation in Video
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2025)
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2025)
Residual Gaussian Splatting for Ultra Sparse-View CBCT Reconstruction
von: Lin, Jian, et al.
Veröffentlicht: (2026)
von: Lin, Jian, et al.
Veröffentlicht: (2026)
Robust Drone-View Geo-Localization via Content-Viewpoint Disentanglement
von: Li, Ke, et al.
Veröffentlicht: (2025)
von: Li, Ke, et al.
Veröffentlicht: (2025)
FIction: 4D Future Interaction Prediction from Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
Seeing the Arrow of Time in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
von: Adebi, Daniel, et al.
Veröffentlicht: (2025)
von: Adebi, Daniel, et al.
Veröffentlicht: (2025)
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
von: Baid, Ami, et al.
Veröffentlicht: (2026)
von: Baid, Ami, et al.
Veröffentlicht: (2026)
Improving Viewpoint-Invariance and Temporal Consistency for Action Detection
von: Porto, Yannick, et al.
Veröffentlicht: (2026)
von: Porto, Yannick, et al.
Veröffentlicht: (2026)
RECIPE: Procedural Planning via Grounding in Instructional Video
von: Seminara, Luigi, et al.
Veröffentlicht: (2026)
von: Seminara, Luigi, et al.
Veröffentlicht: (2026)
Boosting Self-Supervision for Single-View Scene Completion via Knowledge Distillation
von: Han, Keonhee, et al.
Veröffentlicht: (2024)
von: Han, Keonhee, et al.
Veröffentlicht: (2024)
SkillSight: Efficient First-Person Skill Assessment with Gaze
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
EgoExo-WM: Unlocking Exo Video for Ego World Models
von: Tran, Danny, et al.
Veröffentlicht: (2026)
von: Tran, Danny, et al.
Veröffentlicht: (2026)
Progress-Aware Video Frame Captioning
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
SportSkills: Physical Skill Learning from Sports Instructional Videos
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
ViewBridge:Revisiting Cross-View Localization from Image Matching
von: Xia, Panwang, et al.
Veröffentlicht: (2025)
von: Xia, Panwang, et al.
Veröffentlicht: (2025)
VITED: Video Temporal Evidence Distillation
von: Lu, Yujie, et al.
Veröffentlicht: (2025)
von: Lu, Yujie, et al.
Veröffentlicht: (2025)
UniABG: Unified Adversarial View Bridging and Graph Correspondence for Unsupervised Cross-View Geo-Localization
von: Chen, Cuiqun, et al.
Veröffentlicht: (2025)
von: Chen, Cuiqun, et al.
Veröffentlicht: (2025)
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
von: Jung, Minjoon, et al.
Veröffentlicht: (2026)
von: Jung, Minjoon, et al.
Veröffentlicht: (2026)
Multi-Level Embedding and Alignment Network with Consistency and Invariance Learning for Cross-View Geo-Localization
von: Chen, Zhongwei, et al.
Veröffentlicht: (2024)
von: Chen, Zhongwei, et al.
Veröffentlicht: (2024)
AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos
von: Luo, Mi, et al.
Veröffentlicht: (2024)
von: Luo, Mi, et al.
Veröffentlicht: (2024)
Detours for Navigating Instructional Videos
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation
von: Teng, Wenbin, et al.
Veröffentlicht: (2025)
von: Teng, Wenbin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2026) -
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2024) -
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024) -
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
von: Xue, Zihui, et al.
Veröffentlicht: (2024) -
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)