Toward Better SSIM Loss for Unsupervised Monocular Depth Estimation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cao, Yijun, Luo, Fuya, Li, Yongjie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908394488594432
author Cao, Yijun
Luo, Fuya
Li, Yongjie
author_facet Cao, Yijun
Luo, Fuya
Li, Yongjie
contents Unsupervised monocular depth learning generally relies on the photometric relation among temporally adjacent images. Most of previous works use both mean absolute error (MAE) and structure similarity index measure (SSIM) with conventional form as training loss. However, they ignore the effect of different components in the SSIM function and the corresponding hyperparameters on the training. To address these issues, this work proposes a new form of SSIM. Compared with original SSIM function, the proposed new form uses addition rather than multiplication to combine the luminance, contrast, and structural similarity related components in SSIM. The loss function constructed with this scheme helps result in smoother gradients and achieve higher performance on unsupervised depth estimation. We conduct extensive experiments to determine the relatively optimal combination of parameters for our new SSIM. Based on the popular MonoDepth approach, the optimized SSIM loss function can remarkably outperform the baseline on the KITTI-2015 outdoor dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04758
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toward Better SSIM Loss for Unsupervised Monocular Depth Estimation
Cao, Yijun
Luo, Fuya
Li, Yongjie
Computer Vision and Pattern Recognition
Unsupervised monocular depth learning generally relies on the photometric relation among temporally adjacent images. Most of previous works use both mean absolute error (MAE) and structure similarity index measure (SSIM) with conventional form as training loss. However, they ignore the effect of different components in the SSIM function and the corresponding hyperparameters on the training. To address these issues, this work proposes a new form of SSIM. Compared with original SSIM function, the proposed new form uses addition rather than multiplication to combine the luminance, contrast, and structural similarity related components in SSIM. The loss function constructed with this scheme helps result in smoother gradients and achieve higher performance on unsupervised depth estimation. We conduct extensive experiments to determine the relatively optimal combination of parameters for our new SSIM. Based on the popular MonoDepth approach, the optimized SSIM loss function can remarkably outperform the baseline on the KITTI-2015 outdoor dataset.
title Toward Better SSIM Loss for Unsupervised Monocular Depth Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.04758