ObitoNet: Multimodal High-Resolution Point Cloud Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thapliyal, Apoorv, Lanka, Vinay, Baskaran, Swathi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915080722972672
author Thapliyal, Apoorv
Lanka, Vinay
Baskaran, Swathi
author_facet Thapliyal, Apoorv
Lanka, Vinay
Baskaran, Swathi
contents ObitoNet employs a Cross Attention mechanism to integrate multimodal inputs, where Vision Transformers (ViT) extract semantic features from images and a point cloud tokenizer processes geometric information using Farthest Point Sampling (FPS) and K Nearest Neighbors (KNN) for spatial structure capture. The learned multimodal features are fed into a transformer-based decoder for high-resolution point cloud reconstruction. This approach leverages the complementary strengths of both modalities rich image features and precise geometric details ensuring robust point cloud generation even in challenging conditions such as sparse or noisy data.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18775
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ObitoNet: Multimodal High-Resolution Point Cloud Reconstruction
Thapliyal, Apoorv
Lanka, Vinay
Baskaran, Swathi
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
ObitoNet employs a Cross Attention mechanism to integrate multimodal inputs, where Vision Transformers (ViT) extract semantic features from images and a point cloud tokenizer processes geometric information using Farthest Point Sampling (FPS) and K Nearest Neighbors (KNN) for spatial structure capture. The learned multimodal features are fed into a transformer-based decoder for high-resolution point cloud reconstruction. This approach leverages the complementary strengths of both modalities rich image features and precise geometric details ensuring robust point cloud generation even in challenging conditions such as sparse or noisy data.
title ObitoNet: Multimodal High-Resolution Point Cloud Reconstruction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.18775