M2H-MX: Multi-Task Semantic and Geometric Perception for Real-Time Monocular 3D Scene Graph Construction
Fuente:
arXiv
Saved in:
| Main Authors: | Udugama, U. V. B. L., Vosselman, George, Nex, Francesco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mono-Hydra++: Real-Time Monocular Scene Graph Construction with Multi-Task Learning for 3D Indoor Mapping
by: Udugama, U. V. B. L., et al.
Published: (2026)
by: Udugama, U. V. B. L., et al.
Published: (2026)
M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
by: Udugama, U. V. B. L, et al.
Published: (2025)
by: Udugama, U. V. B. L, et al.
Published: (2025)
Mono-hydra: Real-time 3D scene graph construction from monocular camera input with IMU
by: Udugama, U. V. B. L., et al.
Published: (2023)
by: Udugama, U. V. B. L., et al.
Published: (2023)
A Comparison of Multi-View Stereo Methods for Photogrammetric 3D Reconstruction: From Traditional to Learning-Based Approaches
by: Li, Yawen, et al.
Published: (2026)
by: Li, Yawen, et al.
Published: (2026)
Incremental Semantics-Aided Meshing from LiDAR-Inertial Odometry and RGB Direct Label Transfer
by: Affan, Muhammad, et al.
Published: (2026)
by: Affan, Muhammad, et al.
Published: (2026)
ZeD-MAP: Bundle Adjustment Guided Zero-Shot Depth Maps for Real-Time Aerial Imaging
by: Iz, Selim Ahmet, et al.
Published: (2026)
by: Iz, Selim Ahmet, et al.
Published: (2026)
Real-time Monocular 2D and 3D Perception of Endoluminal Scenes for Controlling Flexible Robotic Endoscopic Instruments
by: Wei, Ruofeng, et al.
Published: (2026)
by: Wei, Ruofeng, et al.
Published: (2026)
SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and Its Downstream Tasks
by: Xie, Yaxu, et al.
Published: (2024)
by: Xie, Yaxu, et al.
Published: (2024)
Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion
by: Liang, Li, et al.
Published: (2025)
by: Liang, Li, et al.
Published: (2025)
Real-time 3D Semantic Scene Perception for Egocentric Robots with Binocular Vision
by: Nguyen, K., et al.
Published: (2024)
by: Nguyen, K., et al.
Published: (2024)
Real-Time Monocular Scene Analysis for UAV in Outdoor Environments
by: AlaaEldin, Yara
Published: (2026)
by: AlaaEldin, Yara
Published: (2026)
Real-Time Bundle Adjustment for Ultra-High-Resolution UAV Imagery Using Adaptive Patch-Based Feature Tracking
by: Iz, Selim Ahmet, et al.
Published: (2025)
by: Iz, Selim Ahmet, et al.
Published: (2025)
VexNex
by: VexNex
Published: (2026)
by: VexNex
Published: (2026)
Toward a Real-Time Framework for Accurate Monocular 3D Human Pose Estimation with Geometric Priors
by: Adjel, Mohamed
Published: (2025)
by: Adjel, Mohamed
Published: (2025)
Unleashing Semantic and Geometric Priors for 3D Scene Completion
by: Chen, Shiyuan, et al.
Published: (2025)
by: Chen, Shiyuan, et al.
Published: (2025)
MGNet: Monocular Geometric Scene Understanding for Autonomous Driving
by: Schön, Markus, et al.
Published: (2022)
by: Schön, Markus, et al.
Published: (2022)
MGNiceNet: Unified Monocular Geometric Scene Understanding
by: Schön, Markus, et al.
Published: (2024)
by: Schön, Markus, et al.
Published: (2024)
FROSS: Faster-than-Real-Time Online 3D Semantic Scene Graph Generation from RGB-D Images
by: Hou, Hao-Yu, et al.
Published: (2025)
by: Hou, Hao-Yu, et al.
Published: (2025)
Real-Time 3D Occupancy Prediction via Geometric-Semantic Disentanglement
by: He, Yulin, et al.
Published: (2024)
by: He, Yulin, et al.
Published: (2024)
Clio: Real-time Task-Driven Open-Set 3D Scene Graphs
by: Maggio, Dominic, et al.
Published: (2024)
by: Maggio, Dominic, et al.
Published: (2024)
Transfer Learning from Simulated to Real Scenes for Monocular 3D Object Detection
by: Mohamed, Sondos, et al.
Published: (2024)
by: Mohamed, Sondos, et al.
Published: (2024)
GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
by: Saxena, Saumya, et al.
Published: (2024)
by: Saxena, Saumya, et al.
Published: (2024)
SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos
by: Liu, Yuzheng, et al.
Published: (2024)
by: Liu, Yuzheng, et al.
Published: (2024)
PolyR-CNN: R-CNN for end-to-end polygonal building outline extraction
by: Jiao, Weiqin, et al.
Published: (2024)
by: Jiao, Weiqin, et al.
Published: (2024)
LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
by: Meier, Johannes, et al.
Published: (2025)
by: Meier, Johannes, et al.
Published: (2025)
Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model
by: Wu, Haiyang, et al.
Published: (2026)
by: Wu, Haiyang, et al.
Published: (2026)
Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding
by: Longo, Antonello, et al.
Published: (2025)
by: Longo, Antonello, et al.
Published: (2025)
REACT: Real-time Efficient Attribute Clustering and Transfer for Updatable 3D Scene Graph
by: Nguyen, Phuoc, et al.
Published: (2025)
by: Nguyen, Phuoc, et al.
Published: (2025)
GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
by: Ruiz, Antonio, et al.
Published: (2025)
by: Ruiz, Antonio, et al.
Published: (2025)
Planner3D: LLM-enhanced graph prior meets 3D indoor scene explicit regularization
by: Wei, Yao, et al.
Published: (2024)
by: Wei, Yao, et al.
Published: (2024)
Multimodal Rationales for Explainable Visual Question Answering
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Scale-wise Bidirectional Alignment Network for Referring Remote Sensing Image Segmentation
by: Li, Kun, et al.
Published: (2025)
by: Li, Kun, et al.
Published: (2025)
Semantic Flow: Learning Semantic Field of Dynamic Scenes from Monocular Videos
by: Tian, Fengrui, et al.
Published: (2024)
by: Tian, Fengrui, et al.
Published: (2024)
DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
by: Chu, Wen-Hsuan, et al.
Published: (2024)
by: Chu, Wen-Hsuan, et al.
Published: (2024)
Monocular Semantic Scene Completion via Masked Recurrent Networks
by: Wang, Xuzhi, et al.
Published: (2025)
by: Wang, Xuzhi, et al.
Published: (2025)
Task and Motion Planning in Hierarchical 3D Scene Graphs
by: Ray, Aaron, et al.
Published: (2024)
by: Ray, Aaron, et al.
Published: (2024)
DepthSSC: Monocular 3D Semantic Scene Completion via Depth-Spatial Alignment and Voxel Adaptation
by: Yao, Jiawei, et al.
Published: (2023)
by: Yao, Jiawei, et al.
Published: (2023)
ET-Former: Efficient Triplane Deformable Attention for 3D Semantic Scene Completion From Monocular Camera
by: Liang, Jing, et al.
Published: (2024)
by: Liang, Jing, et al.
Published: (2024)
Enhancing Monocular 3D Scene Completion with Diffusion Model
by: Song, Changlin, et al.
Published: (2025)
by: Song, Changlin, et al.
Published: (2025)
SMPISD-MTPNet: Scene Semantic Prior-Assisted Infrared Ship Detection Using Multi-Task Perception Networks
by: Hu, Chen, et al.
Published: (2024)
by: Hu, Chen, et al.
Published: (2024)
Similar Items
-
Mono-Hydra++: Real-Time Monocular Scene Graph Construction with Multi-Task Learning for 3D Indoor Mapping
by: Udugama, U. V. B. L., et al.
Published: (2026) -
M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
by: Udugama, U. V. B. L, et al.
Published: (2025) -
Mono-hydra: Real-time 3D scene graph construction from monocular camera input with IMU
by: Udugama, U. V. B. L., et al.
Published: (2023) -
A Comparison of Multi-View Stereo Methods for Photogrammetric 3D Reconstruction: From Traditional to Learning-Based Approaches
by: Li, Yawen, et al.
Published: (2026) -
Incremental Semantics-Aided Meshing from LiDAR-Inertial Odometry and RGB Direct Label Transfer
by: Affan, Muhammad, et al.
Published: (2026)