BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Hosseinzadeh, Mehdi, Reid, Ian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-BEV: Zero-shot Projection of Any First-Person Modality to BEV Maps
by: Monaci, Gianluca, et al.
Published: (2024)
by: Monaci, Gianluca, et al.
Published: (2024)
UniBEV: Multi-modal 3D Object Detection with Uniform BEV Encoders for Robustness against Missing Sensor Modalities
by: Wang, Shiming, et al.
Published: (2023)
by: Wang, Shiming, et al.
Published: (2023)
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
by: Wang, Wenze, et al.
Published: (2026)
by: Wang, Wenze, et al.
Published: (2026)
SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment
by: Singh, Binod, et al.
Published: (2025)
by: Singh, Binod, et al.
Published: (2025)
DiffSemanticFusion: Semantic Raster BEV Fusion for Autonomous Driving via Online HD Map Diffusion
by: Sun, Zhigang, et al.
Published: (2025)
by: Sun, Zhigang, et al.
Published: (2025)
Bridging Perspectives: Foundation Model Guided BEV Maps for 3D Object Detection and Tracking
by: Käppeler, Markus, et al.
Published: (2025)
by: Käppeler, Markus, et al.
Published: (2025)
BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images
by: Skuddis, David, et al.
Published: (2026)
by: Skuddis, David, et al.
Published: (2026)
BEV-Patch-PF: Particle Filtering with BEV-Aerial Feature Matching for Off-Road Geo-Localization
by: Lee, Dongmyeong, et al.
Published: (2025)
by: Lee, Dongmyeong, et al.
Published: (2025)
Benchmarking Multi-View BEV Object Detection with Mixed Pinhole and Fisheye Cameras
by: Liu, Xiangzhong, et al.
Published: (2026)
by: Liu, Xiangzhong, et al.
Published: (2026)
LetsMap: Unsupervised Representation Learning for Semantic BEV Mapping
by: Gosala, Nikhil, et al.
Published: (2024)
by: Gosala, Nikhil, et al.
Published: (2024)
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
by: Man, Yunze, et al.
Published: (2023)
by: Man, Yunze, et al.
Published: (2023)
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
by: Podgorski, Stefan, et al.
Published: (2025)
by: Podgorski, Stefan, et al.
Published: (2025)
Incremental Joint Learning of Depth, Pose and Implicit Scene Representation on Monocular Camera in Large-scale Scenes
by: Deng, Tianchen, et al.
Published: (2024)
by: Deng, Tianchen, et al.
Published: (2024)
UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation
by: Jiang, Zhaodong, et al.
Published: (2025)
by: Jiang, Zhaodong, et al.
Published: (2025)
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
by: Hosseinzadeh, Mehdi, et al.
Published: (2026)
by: Hosseinzadeh, Mehdi, et al.
Published: (2026)
DiffusionNOCS: Managing Symmetry and Uncertainty in Sim2Real Multi-Modal Category-level Pose Estimation
by: Ikeda, Takuya, et al.
Published: (2024)
by: Ikeda, Takuya, et al.
Published: (2024)
OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGB
by: Lin, Yunzhi, et al.
Published: (2024)
by: Lin, Yunzhi, et al.
Published: (2024)
Alignment Scores: Robust Metrics for Multiview Pose Accuracy Evaluation
by: Lee, Seong Hun, et al.
Published: (2024)
by: Lee, Seong Hun, et al.
Published: (2024)
Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
by: Li, Xun, et al.
Published: (2025)
by: Li, Xun, et al.
Published: (2025)
BEVCar: Camera-Radar Fusion for BEV Map and Object Segmentation
by: Schramm, Jonas, et al.
Published: (2024)
by: Schramm, Jonas, et al.
Published: (2024)
RFM-Pose:Reinforcement-Guided Flow Matching for Fast Category-Level 6D Pose Estimation
by: He, Diya, et al.
Published: (2026)
by: He, Diya, et al.
Published: (2026)
SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and Its Downstream Tasks
by: Xie, Yaxu, et al.
Published: (2024)
by: Xie, Yaxu, et al.
Published: (2024)
MASSTAR: A Multi-Modal and Large-Scale Scene Dataset with a Versatile Toolchain for Surface Prediction and Completion
by: Zheng, Guiyong, et al.
Published: (2024)
by: Zheng, Guiyong, et al.
Published: (2024)
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
by: Garg, Sourav, et al.
Published: (2024)
by: Garg, Sourav, et al.
Published: (2024)
GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection
by: Song, Ziying, et al.
Published: (2024)
by: Song, Ziying, et al.
Published: (2024)
Object Pose Estimation through Dexterous Touch
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2025)
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2025)
APR-Transformer: Initial Pose Estimation for Localization in Complex Environments through Absolute Pose Regression
by: Ravuri, Srinivas, et al.
Published: (2025)
by: Ravuri, Srinivas, et al.
Published: (2025)
BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving
by: Winter, Katharina, et al.
Published: (2025)
by: Winter, Katharina, et al.
Published: (2025)
OccCylindrical: Multi-Modal Fusion with Cylindrical Representation for 3D Semantic Occupancy Prediction
by: Ming, Zhenxing, et al.
Published: (2025)
by: Ming, Zhenxing, et al.
Published: (2025)
DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
by: Gao, Yu, et al.
Published: (2025)
by: Gao, Yu, et al.
Published: (2025)
Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation
by: Li, Haosheng, et al.
Published: (2024)
by: Li, Haosheng, et al.
Published: (2024)
Simulation-Ready Cluttered Scene Estimation via Physics-aware Joint Shape and Pose Optimization
by: Huang, Wei-Cheng, et al.
Published: (2026)
by: Huang, Wei-Cheng, et al.
Published: (2026)
Label-efficient Semantic Scene Completion with Scribble Annotations
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Multi-view Pose Fusion for Occlusion-Aware 3D Human Pose Estimation
by: Bragagnolo, Laura, et al.
Published: (2024)
by: Bragagnolo, Laura, et al.
Published: (2024)
VLMFusionOcc3D: VLM Assisted Multi-Modal 3D Semantic Occupancy Prediction
by: Doruk, A. Enes, et al.
Published: (2026)
by: Doruk, A. Enes, et al.
Published: (2026)
Collaborative Representation Learning for Alignment of Tactile, Language, and Vision Modalities
by: Zhou, Yiyun, et al.
Published: (2025)
by: Zhou, Yiyun, et al.
Published: (2025)
Pose-Guided Residual Refinement for Interpretable Text-to-Motion Generation and Editing
by: Jeong, Sukhyun, et al.
Published: (2025)
by: Jeong, Sukhyun, et al.
Published: (2025)
GeoWorld: Geometric World Models
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
Latent Action Pretraining Through World Modeling
by: Tharwat, Bahey, et al.
Published: (2025)
by: Tharwat, Bahey, et al.
Published: (2025)
Watch Your STEPP: Semantic Traversability Estimation using Pose Projected Features
by: Ægidius, Sebastian, et al.
Published: (2025)
by: Ægidius, Sebastian, et al.
Published: (2025)
Similar Items
-
Zero-BEV: Zero-shot Projection of Any First-Person Modality to BEV Maps
by: Monaci, Gianluca, et al.
Published: (2024) -
UniBEV: Multi-modal 3D Object Detection with Uniform BEV Encoders for Robustness against Missing Sensor Modalities
by: Wang, Shiming, et al.
Published: (2023) -
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
by: Wang, Wenze, et al.
Published: (2026) -
SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment
by: Singh, Binod, et al.
Published: (2025) -
DiffSemanticFusion: Semantic Raster BEV Fusion for Autonomous Driving via Online HD Map Diffusion
by: Sun, Zhigang, et al.
Published: (2025)