Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
Fuente:
arXiv
Saved in:
| Main Authors: | Adebi, Daniel, Majumder, Sagnik, Grauman, Kristen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
by: Majumder, Sagnik, et al.
Published: (2023)
by: Majumder, Sagnik, et al.
Published: (2023)
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
by: Somayazulu, Arjun, et al.
Published: (2024)
by: Somayazulu, Arjun, et al.
Published: (2024)
MistExit: Learning to Exit for Early Mistake Detection in Procedural Videos
by: Majumder, Sagnik, et al.
Published: (2026)
by: Majumder, Sagnik, et al.
Published: (2026)
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
by: Majumder, Sagnik, et al.
Published: (2024)
by: Majumder, Sagnik, et al.
Published: (2024)
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
by: Baid, Ami, et al.
Published: (2026)
by: Baid, Ami, et al.
Published: (2026)
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos
by: Majumder, Sagnik, et al.
Published: (2024)
by: Majumder, Sagnik, et al.
Published: (2024)
Few-View Object Reconstruction with Unknown Categories and Camera Poses
by: Jiang, Hanwen, et al.
Published: (2022)
by: Jiang, Hanwen, et al.
Published: (2022)
Mash, Spread, Slice! Learning to Manipulate Object States via Visual Spatial Progress
by: Mandikal, Priyanka, et al.
Published: (2025)
by: Mandikal, Priyanka, et al.
Published: (2025)
Materialistic RIR: Material Conditioned Realistic RIR Generation
by: Saad, Mahnoor Fatima, et al.
Published: (2026)
by: Saad, Mahnoor Fatima, et al.
Published: (2026)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
by: An, Joungbin, et al.
Published: (2025)
by: An, Joungbin, et al.
Published: (2025)
Learning Skill-Attributes for Transferable Assessment in Video
by: Ashutosh, Kumar, et al.
Published: (2025)
by: Ashutosh, Kumar, et al.
Published: (2025)
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
by: Somayazulu, Arjun, et al.
Published: (2026)
by: Somayazulu, Arjun, et al.
Published: (2026)
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
by: An, Joungbin, et al.
Published: (2026)
by: An, Joungbin, et al.
Published: (2026)
Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos
by: Sun, Shuo, et al.
Published: (2026)
by: Sun, Shuo, et al.
Published: (2026)
FIction: 4D Future Interaction Prediction from Video
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
Learning Object State Changes in Videos: An Open-World Perspective
by: Xue, Zihui, et al.
Published: (2023)
by: Xue, Zihui, et al.
Published: (2023)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
Seeing without Pixels: Perception from Camera Trajectories
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
by: Wu, Chi Hsuan, et al.
Published: (2025)
by: Wu, Chi Hsuan, et al.
Published: (2025)
EgoExo-WM: Unlocking Exo Video for Ego World Models
by: Tran, Danny, et al.
Published: (2026)
by: Tran, Danny, et al.
Published: (2026)
SportSkills: Physical Skill Learning from Sports Instructional Videos
by: Ashutosh, Kumar, et al.
Published: (2026)
by: Ashutosh, Kumar, et al.
Published: (2026)
Detours for Navigating Instructional Videos
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos
by: Luo, Mi, et al.
Published: (2024)
by: Luo, Mi, et al.
Published: (2024)
WildPose: A Unified Framework for Robust Pose Estimation in the Wild
by: Zheng, Jianhao, et al.
Published: (2026)
by: Zheng, Jianhao, et al.
Published: (2026)
ExpertAF: Expert Actionable Feedback from Video
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
Personal Visual Context Learning in Large Multimodal Models
by: Xue, Zihui, et al.
Published: (2026)
by: Xue, Zihui, et al.
Published: (2026)
SPOC: Spatially-Progressing Object State Change Segmentation in Video
by: Mandikal, Priyanka, et al.
Published: (2025)
by: Mandikal, Priyanka, et al.
Published: (2025)
Seeing the Arrow of Time in Large Multimodal Models
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
by: Luo, Mi, et al.
Published: (2025)
by: Luo, Mi, et al.
Published: (2025)
Free-DyGS: Camera-Pose-Free Scene Reconstruction for Dynamic Surgical Videos with Gaussian Splatting
by: Li, Qian, et al.
Published: (2024)
by: Li, Qian, et al.
Published: (2024)
SkillSight: Efficient First-Person Skill Assessment with Gaze
by: Wu, Chi Hsuan, et al.
Published: (2025)
by: Wu, Chi Hsuan, et al.
Published: (2025)
PoseFM: Relative Camera Pose Estimation Through Flow Matching
by: Kuczkowski, Dominik, et al.
Published: (2026)
by: Kuczkowski, Dominik, et al.
Published: (2026)
Semi-Supervised Unconstrained Head Pose Estimation in the Wild
by: Zhou, Huayi, et al.
Published: (2024)
by: Zhou, Huayi, et al.
Published: (2024)
AudioScenic: Audio-Driven Video Scene Editing
by: Shen, Kaixin, et al.
Published: (2024)
by: Shen, Kaixin, et al.
Published: (2024)
CLHOP: Combined Audio-Video Learning for Horse 3D Pose and Shape Estimation
by: Li, Ci, et al.
Published: (2024)
by: Li, Ci, et al.
Published: (2024)
AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual Videos
by: Wimbauer, Felix, et al.
Published: (2025)
by: Wimbauer, Felix, et al.
Published: (2025)
Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
Similar Items
-
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
by: Majumder, Sagnik, et al.
Published: (2023) -
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
by: Somayazulu, Arjun, et al.
Published: (2024) -
MistExit: Learning to Exit for Early Mistake Detection in Procedural Videos
by: Majumder, Sagnik, et al.
Published: (2026) -
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
by: Majumder, Sagnik, et al.
Published: (2024) -
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
by: Baid, Ami, et al.
Published: (2026)