Mitigating Perspective Distortion-induced Shape Ambiguity in Image Crops

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Prakash, Aditya, Gupta, Arjun, Gupta, Saurabh
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916406454386688
author Prakash, Aditya
Gupta, Arjun
Gupta, Saurabh
author_facet Prakash, Aditya
Gupta, Arjun
Gupta, Saurabh
contents Objects undergo varying amounts of perspective distortion as they move across a camera's field of view. Models for predicting 3D from a single image often work with crops around the object of interest and ignore the location of the object in the camera's field of view. We note that ignoring this location information further exaggerates the inherent ambiguity in making 3D inferences from 2D images and can prevent models from even fitting to the training data. To mitigate this ambiguity, we propose Intrinsics-Aware Positional Encoding (KPE), which incorporates information about the location of crops in the image and camera intrinsics. Experiments on three popular 3D-from-a-single-image benchmarks: depth prediction on NYU, 3D object detection on KITTI & nuScenes, and predicting 3D shapes of articulated objects on ARCTIC, show the benefits of KPE.
format Preprint
id arxiv_https___arxiv_org_abs_2312_06594
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Mitigating Perspective Distortion-induced Shape Ambiguity in Image Crops
Prakash, Aditya
Gupta, Arjun
Gupta, Saurabh
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Objects undergo varying amounts of perspective distortion as they move across a camera's field of view. Models for predicting 3D from a single image often work with crops around the object of interest and ignore the location of the object in the camera's field of view. We note that ignoring this location information further exaggerates the inherent ambiguity in making 3D inferences from 2D images and can prevent models from even fitting to the training data. To mitigate this ambiguity, we propose Intrinsics-Aware Positional Encoding (KPE), which incorporates information about the location of crops in the image and camera intrinsics. Experiments on three popular 3D-from-a-single-image benchmarks: depth prediction on NYU, 3D object detection on KITTI & nuScenes, and predicting 3D shapes of articulated objects on ARCTIC, show the benefits of KPE.
title Mitigating Perspective Distortion-induced Shape Ambiguity in Image Crops
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.06594