M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Udugama, U. V. B. L, Vosselman, George, Nex, Francesco |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mono-Hydra++: Real-Time Monocular Scene Graph Construction with Multi-Task Learning for 3D Indoor Mapping
by: Udugama, U. V. B. L., et al.
Published: (2026)
by: Udugama, U. V. B. L., et al.
Published: (2026)
M2H-MX: Multi-Task Semantic and Geometric Perception for Real-Time Monocular 3D Scene Graph Construction
by: Udugama, U. V. B. L., et al.
Published: (2026)
by: Udugama, U. V. B. L., et al.
Published: (2026)
Mono-hydra: Real-time 3D scene graph construction from monocular camera input with IMU
by: Udugama, U. V. B. L., et al.
Published: (2023)
by: Udugama, U. V. B. L., et al.
Published: (2023)
A Comparison of Multi-View Stereo Methods for Photogrammetric 3D Reconstruction: From Traditional to Learning-Based Approaches
by: Li, Yawen, et al.
Published: (2026)
by: Li, Yawen, et al.
Published: (2026)
Multi-Task Learning for Robot Perception with Imbalanced Data
by: Erkent, Ozgur
Published: (2026)
by: Erkent, Ozgur
Published: (2026)
Incremental Semantics-Aided Meshing from LiDAR-Inertial Odometry and RGB Direct Label Transfer
by: Affan, Muhammad, et al.
Published: (2026)
by: Affan, Muhammad, et al.
Published: (2026)
EagleVision: A Multi-Task Benchmark for Cross-Domain Perception in High-Speed Autonomous Racing
by: Yagudin, Zakhar, et al.
Published: (2026)
by: Yagudin, Zakhar, et al.
Published: (2026)
EvidMTL: Evidential Multi-Task Learning for Uncertainty-Aware Semantic Surface Mapping from Monocular RGB Images
by: Menon, Rohit, et al.
Published: (2025)
by: Menon, Rohit, et al.
Published: (2025)
Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model
by: Wu, Haiyang, et al.
Published: (2026)
by: Wu, Haiyang, et al.
Published: (2026)
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
by: Man, Yunze, et al.
Published: (2023)
by: Man, Yunze, et al.
Published: (2023)
Residual Vector Quantization For Communication-Efficient Multi-Agent Perception
by: Shenkut, Dereje, et al.
Published: (2025)
by: Shenkut, Dereje, et al.
Published: (2025)
Efficient Multi-Task Scene Analysis with RGB-D Transformers
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023)
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023)
MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
by: Hao, Jinkun, et al.
Published: (2025)
by: Hao, Jinkun, et al.
Published: (2025)
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
by: Qi, Yu, et al.
Published: (2025)
by: Qi, Yu, et al.
Published: (2025)
Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation
by: Ma, Teli, et al.
Published: (2024)
by: Ma, Teli, et al.
Published: (2024)
ZeD-MAP: Bundle Adjustment Guided Zero-Shot Depth Maps for Real-Time Aerial Imaging
by: Iz, Selim Ahmet, et al.
Published: (2026)
by: Iz, Selim Ahmet, et al.
Published: (2026)
DiFuse-Net: RGB and Dual-Pixel Depth Estimation using Window Bi-directional Parallax Attention and Cross-modal Transfer Learning
by: Swami, Kunal, et al.
Published: (2025)
by: Swami, Kunal, et al.
Published: (2025)
SWA-SOP: Spatially-aware Window Attention for Semantic Occupancy Prediction in Autonomous Driving
by: Cao, Helin, et al.
Published: (2025)
by: Cao, Helin, et al.
Published: (2025)
LiDAR-BEVMTN: Real-Time LiDAR Bird's-Eye View Multi-Task Perception Network for Autonomous Driving
by: Mohapatra, Sambit, et al.
Published: (2023)
by: Mohapatra, Sambit, et al.
Published: (2023)
STAMP: Scalable Task And Model-agnostic Collaborative Perception
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception
by: Vosshans, Marcel, et al.
Published: (2025)
by: Vosshans, Marcel, et al.
Published: (2025)
Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations
by: Goko, Miyu, et al.
Published: (2024)
by: Goko, Miyu, et al.
Published: (2024)
Motion Consistency Loss for Monocular Visual Odometry with Attention-Based Deep Learning
by: Françani, André O., et al.
Published: (2024)
by: Françani, André O., et al.
Published: (2024)
EGSA-PT:Edge-Guided Spatial Attention with Progressive Training for Monocular Depth Estimation and Segmentation of Transparent Objects
by: Omotara, Gbenga, et al.
Published: (2025)
by: Omotara, Gbenga, et al.
Published: (2025)
Next-Future: Sample-Efficient Policy Learning for Robotic-Arm Tasks
by: Özgür, Fikrican, et al.
Published: (2025)
by: Özgür, Fikrican, et al.
Published: (2025)
RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion
by: Pataki, Zador, et al.
Published: (2025)
by: Pataki, Zador, et al.
Published: (2025)
Task-Oriented Human Grasp Synthesis via Context- and Task-Aware Diffusers
by: Liu, An-Lun, et al.
Published: (2025)
by: Liu, An-Lun, et al.
Published: (2025)
A Survey on Deep Multi-Task Learning in Connected Autonomous Vehicles
by: Wang, Jiayuan, et al.
Published: (2025)
by: Wang, Jiayuan, et al.
Published: (2025)
DNAct: Diffusion Guided Multi-Task 3D Policy Learning
by: Yan, Ge, et al.
Published: (2024)
by: Yan, Ge, et al.
Published: (2024)
AsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory
by: Ma, Lianjie, et al.
Published: (2026)
by: Ma, Lianjie, et al.
Published: (2026)
Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation
by: Zhang, Xitie, et al.
Published: (2026)
by: Zhang, Xitie, et al.
Published: (2026)
Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks
by: Łucki, Jakub, et al.
Published: (2025)
by: Łucki, Jakub, et al.
Published: (2025)
Monocular Visual Place Recognition in LiDAR Maps via Cross-Modal State Space Model and Multi-View Matching
by: Yao, Gongxin, et al.
Published: (2024)
by: Yao, Gongxin, et al.
Published: (2024)
Masked Depth Modeling for Spatial Perception
by: Tan, Bin, et al.
Published: (2026)
by: Tan, Bin, et al.
Published: (2026)
Enhancing Video-Based Robot Failure Detection Using Task Knowledge
by: Thoduka, Santosh, et al.
Published: (2025)
by: Thoduka, Santosh, et al.
Published: (2025)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
by: Shi, Jiaqi, et al.
Published: (2026)
by: Shi, Jiaqi, et al.
Published: (2026)
Towards Dense and Accurate Radar Perception Via Efficient Cross-Modal Diffusion Model
by: Zhang, Ruibin, et al.
Published: (2024)
by: Zhang, Ruibin, et al.
Published: (2024)
VisLanding: Monocular 3D Perception for UAV Safe Landing via Depth-Normal Synergy
by: Tan, Zhuoyue, et al.
Published: (2025)
by: Tan, Zhuoyue, et al.
Published: (2025)
VertiFormer: A Data-Efficient Multi-Task Transformer for Off-Road Robot Mobility
by: Nazeri, Mohammad, et al.
Published: (2025)
by: Nazeri, Mohammad, et al.
Published: (2025)
Similar Items
-
Mono-Hydra++: Real-Time Monocular Scene Graph Construction with Multi-Task Learning for 3D Indoor Mapping
by: Udugama, U. V. B. L., et al.
Published: (2026) -
M2H-MX: Multi-Task Semantic and Geometric Perception for Real-Time Monocular 3D Scene Graph Construction
by: Udugama, U. V. B. L., et al.
Published: (2026) -
Mono-hydra: Real-time 3D scene graph construction from monocular camera input with IMU
by: Udugama, U. V. B. L., et al.
Published: (2023) -
A Comparison of Multi-View Stereo Methods for Photogrammetric 3D Reconstruction: From Traditional to Learning-Based Approaches
by: Li, Yawen, et al.
Published: (2026) -
Multi-Task Learning for Robot Perception with Imbalanced Data
by: Erkent, Ozgur
Published: (2026)