Deep Learning Aided Vision System for Planetary Rovers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Relia, Lomash, Singla, Jai G, Amitabh, Dube, Nitant
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914427194834944
author Relia, Lomash
Singla, Jai G
Amitabh
Dube, Nitant
author_facet Relia, Lomash
Singla, Jai G
Amitabh
Dube, Nitant
contents This study presents a vision system for planetary rovers, combining real-time perception with offline terrain reconstruction. The real-time module integrates CLAHE enhanced stereo imagery, YOLOv11n based object detection, and a neural network to estimate object distances. The offline module uses the Depth Anything V2 metric monocular depth estimation model to generate depth maps from captured images, which are fused into dense point clouds using Open3D. Real world distance estimates from the real time pipeline provide reliable metric context alongside the qualitative reconstructions. Evaluation on Chandrayaan 3 NavCam stereo imagery, benchmarked against a CAHV based utility, shows that the neural network achieves a median depth error of 2.26 cm within a 1 to 10 meter range. The object detection model maintains a balanced precision recall tradeoff on grayscale lunar scenes. This architecture offers a scalable, compute-efficient vision solution for autonomous planetary exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26802
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Deep Learning Aided Vision System for Planetary Rovers
Relia, Lomash
Singla, Jai G
Amitabh
Dube, Nitant
Computer Vision and Pattern Recognition
Robotics
Image and Video Processing
This study presents a vision system for planetary rovers, combining real-time perception with offline terrain reconstruction. The real-time module integrates CLAHE enhanced stereo imagery, YOLOv11n based object detection, and a neural network to estimate object distances. The offline module uses the Depth Anything V2 metric monocular depth estimation model to generate depth maps from captured images, which are fused into dense point clouds using Open3D. Real world distance estimates from the real time pipeline provide reliable metric context alongside the qualitative reconstructions. Evaluation on Chandrayaan 3 NavCam stereo imagery, benchmarked against a CAHV based utility, shows that the neural network achieves a median depth error of 2.26 cm within a 1 to 10 meter range. The object detection model maintains a balanced precision recall tradeoff on grayscale lunar scenes. This architecture offers a scalable, compute-efficient vision solution for autonomous planetary exploration.
title Deep Learning Aided Vision System for Planetary Rovers
topic Computer Vision and Pattern Recognition
Robotics
Image and Video Processing
url https://arxiv.org/abs/2603.26802