Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Zhenjiang, Jhong, Dong-You, Wang, Ao, Ruchkin, Ivan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929334642540544
author Mao, Zhenjiang
Jhong, Dong-You
Wang, Ao
Ruchkin, Ivan
author_facet Mao, Zhenjiang
Jhong, Dong-You
Wang, Ao
Ruchkin, Ivan
contents Out-of-distribution (OOD) detection is essential in autonomous driving, to determine when learning-based components encounter unexpected inputs. Traditional detectors typically use encoder models with fixed settings, thus lacking effective human interaction capabilities. With the rise of large foundation models, multimodal inputs offer the possibility of taking human language as a latent representation, thus enabling language-defined OOD detection. In this paper, we use the cosine similarity of image and text representations encoded by the multimodal model CLIP as a new representation to improve the transparency and controllability of latent encodings used for visual anomaly detection. We compare our approach with existing pre-trained encoders that can only produce latent representations that are meaningless from the user's standpoint. Our experiments on realistic driving data show that the language-based latent representation performs better than the traditional representation of the vision encoder and helps improve the detection performance when combined with standard representations.
format Preprint
id arxiv_https___arxiv_org_abs_2405_01691
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving
Mao, Zhenjiang
Jhong, Dong-You
Wang, Ao
Ruchkin, Ivan
Computer Vision and Pattern Recognition
Machine Learning
Robotics
Out-of-distribution (OOD) detection is essential in autonomous driving, to determine when learning-based components encounter unexpected inputs. Traditional detectors typically use encoder models with fixed settings, thus lacking effective human interaction capabilities. With the rise of large foundation models, multimodal inputs offer the possibility of taking human language as a latent representation, thus enabling language-defined OOD detection. In this paper, we use the cosine similarity of image and text representations encoded by the multimodal model CLIP as a new representation to improve the transparency and controllability of latent encodings used for visual anomaly detection. We compare our approach with existing pre-trained encoders that can only produce latent representations that are meaningless from the user's standpoint. Our experiments on realistic driving data show that the language-based latent representation performs better than the traditional representation of the vision encoder and helps improve the detection performance when combined with standard representations.
title Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving
topic Computer Vision and Pattern Recognition
Machine Learning
Robotics
url https://arxiv.org/abs/2405.01691