Motion Aware ViT-based Framework for Monocular 6-DoF Spacecraft Pose Estimation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sosa, Jose, Pineau, Dan, Rathinam, Arunkumar, Shabayek, Abdelrahman, Aouada, Djamila
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911141635031040
author Sosa, Jose
Pineau, Dan
Rathinam, Arunkumar
Shabayek, Abdelrahman
Aouada, Djamila
author_facet Sosa, Jose
Pineau, Dan
Rathinam, Arunkumar
Shabayek, Abdelrahman
Aouada, Djamila
contents Monocular 6-DoF pose estimation plays an important role in multiple spacecraft missions. Most existing pose estimation approaches rely on single images with static keypoint localisation, failing to exploit valuable temporal information inherent to space operations. In this work, we adapt a deep learning framework from human pose estimation to the spacecraft pose estimation domain that integrates motion-aware heatmaps and optical flow to capture motion dynamics. Our approach combines image features from a Vision Transformer (ViT) encoder with motion cues from a pre-trained optical flow model to localise 2D keypoints. Using the estimates, a Perspective-n-Point (PnP) solver recovers 6-DoF poses from known 2D-3D correspondences. We train and evaluate our method on the SPADES-RGB dataset and further assess its generalisation on real and synthetic data from the SPARK-2024 dataset. Overall, our approach demonstrates improved performance over single-image baselines in both 2D keypoint localisation and 6-DoF pose estimation. Furthermore, it shows promising generalisation capabilities when testing on different data distributions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06000
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Motion Aware ViT-based Framework for Monocular 6-DoF Spacecraft Pose Estimation
Sosa, Jose
Pineau, Dan
Rathinam, Arunkumar
Shabayek, Abdelrahman
Aouada, Djamila
Computer Vision and Pattern Recognition
Monocular 6-DoF pose estimation plays an important role in multiple spacecraft missions. Most existing pose estimation approaches rely on single images with static keypoint localisation, failing to exploit valuable temporal information inherent to space operations. In this work, we adapt a deep learning framework from human pose estimation to the spacecraft pose estimation domain that integrates motion-aware heatmaps and optical flow to capture motion dynamics. Our approach combines image features from a Vision Transformer (ViT) encoder with motion cues from a pre-trained optical flow model to localise 2D keypoints. Using the estimates, a Perspective-n-Point (PnP) solver recovers 6-DoF poses from known 2D-3D correspondences. We train and evaluate our method on the SPADES-RGB dataset and further assess its generalisation on real and synthetic data from the SPARK-2024 dataset. Overall, our approach demonstrates improved performance over single-image baselines in both 2D keypoint localisation and 6-DoF pose estimation. Furthermore, it shows promising generalisation capabilities when testing on different data distributions.
title Motion Aware ViT-based Framework for Monocular 6-DoF Spacecraft Pose Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.06000