AsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Lianjie, Li, Yuquan, Jiang, Bingzheng, Zhong, Ziming, Ding, Han, Zhu, Lijun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911505638752256
author Ma, Lianjie
Li, Yuquan
Jiang, Bingzheng
Zhong, Ziming
Ding, Han
Zhu, Lijun
author_facet Ma, Lianjie
Li, Yuquan
Jiang, Bingzheng
Zhong, Ziming
Ding, Han
Zhu, Lijun
contents Foundation-model-based monocular depth estimation offers a viable alternative to active sensors for robot perception, yet its computational cost often prohibits deployment on edge platforms. Existing methods perform independent per-frame inference, wasting the substantial computational redundancy between adjacent viewpoints in continuous robot operation. This paper presents AsyncMDE, an asynchronous depth perception system consisting of a foundation model and a lightweight model that amortizes the foundation model's computational cost over time. The foundation model produces high-quality spatial features in the background, while the lightweight model runs asynchronously in the foreground, fusing cached memory with current observations through complementary fusion, outputting depth estimates, and autoregressively updating the memory. This enables cross-frame feature reuse with bounded accuracy degradation. At a mere 3.83M parameters, it operates at 237 FPS on an RTX 4090, recovering 77% of the accuracy gap to the foundation model while achieving a 25X parameter reduction. Validated across indoor static, dynamic, and synthetic extreme-motion benchmarks, AsyncMDE degrades gracefully between refreshes and achieves 161FPS on a Jetson AGX Orin with TensorRT, clearly demonstrating its feasibility for real-time edge deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10438
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory
Ma, Lianjie
Li, Yuquan
Jiang, Bingzheng
Zhong, Ziming
Ding, Han
Zhu, Lijun
Robotics
Computer Vision and Pattern Recognition
Foundation-model-based monocular depth estimation offers a viable alternative to active sensors for robot perception, yet its computational cost often prohibits deployment on edge platforms. Existing methods perform independent per-frame inference, wasting the substantial computational redundancy between adjacent viewpoints in continuous robot operation. This paper presents AsyncMDE, an asynchronous depth perception system consisting of a foundation model and a lightweight model that amortizes the foundation model's computational cost over time. The foundation model produces high-quality spatial features in the background, while the lightweight model runs asynchronously in the foreground, fusing cached memory with current observations through complementary fusion, outputting depth estimates, and autoregressively updating the memory. This enables cross-frame feature reuse with bounded accuracy degradation. At a mere 3.83M parameters, it operates at 237 FPS on an RTX 4090, recovering 77% of the accuracy gap to the foundation model while achieving a 25X parameter reduction. Validated across indoor static, dynamic, and synthetic extreme-motion benchmarks, AsyncMDE degrades gracefully between refreshes and achieves 161FPS on a Jetson AGX Orin with TensorRT, clearly demonstrating its feasibility for real-time edge deployment.
title AsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.10438