ObitoNet: Multimodal High-Resolution Point Cloud Reconstruction
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915080722972672 |
|---|---|
| author | Thapliyal, Apoorv Lanka, Vinay Baskaran, Swathi |
| author_facet | Thapliyal, Apoorv Lanka, Vinay Baskaran, Swathi |
| contents | ObitoNet employs a Cross Attention mechanism to integrate multimodal inputs, where Vision Transformers (ViT) extract semantic features from images and a point cloud tokenizer processes geometric information using Farthest Point Sampling (FPS) and K Nearest Neighbors (KNN) for spatial structure capture. The learned multimodal features are fed into a transformer-based decoder for high-resolution point cloud reconstruction. This approach leverages the complementary strengths of both modalities rich image features and precise geometric details ensuring robust point cloud generation even in challenging conditions such as sparse or noisy data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_18775 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ObitoNet: Multimodal High-Resolution Point Cloud Reconstruction Thapliyal, Apoorv Lanka, Vinay Baskaran, Swathi Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning ObitoNet employs a Cross Attention mechanism to integrate multimodal inputs, where Vision Transformers (ViT) extract semantic features from images and a point cloud tokenizer processes geometric information using Farthest Point Sampling (FPS) and K Nearest Neighbors (KNN) for spatial structure capture. The learned multimodal features are fed into a transformer-based decoder for high-resolution point cloud reconstruction. This approach leverages the complementary strengths of both modalities rich image features and precise geometric details ensuring robust point cloud generation even in challenging conditions such as sparse or noisy data. |
| title | ObitoNet: Multimodal High-Resolution Point Cloud Reconstruction |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2412.18775 |