TSM-Pose: Topology-Aware Learning with Semantic Mamba for Category-Level Object Pose Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jinshuo, Ma, Bingtao, Su, Junlin, Pan, Guanyuan, Wu, Beining, Yang, Cheng, Lu, Jiaxuan, Yan, Chenggang, Wang, Shuai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911604124155904
author Liu, Jinshuo
Ma, Bingtao
Su, Junlin
Pan, Guanyuan
Wu, Beining
Yang, Cheng
Lu, Jiaxuan
Yan, Chenggang
Wang, Shuai
author_facet Liu, Jinshuo
Ma, Bingtao
Su, Junlin
Pan, Guanyuan
Wu, Beining
Yang, Cheng
Lu, Jiaxuan
Yan, Chenggang
Wang, Shuai
contents Category-level object pose estimation is fundamental for embodied intelligence, yet achieving robust generalization to unseen instances remains challenging. However, existing methods mainly rely on simple feature extraction and aggregation, which struggle to capture category-shared topological structures and conduct semantic keypoint modeling, limiting their generalization. To address these, we propose a \textbf{T}opology-Aware Learning with \textbf{S}emantic \textbf{M}amba for Category-Level \textbf{P}ose Estimation framework (TSM-Pose). Specifically, we introduce a Topology Extractor to capture the global topological representation of the point cloud, which is integrated into local geometry features and enables robust category-level structural representation. Simultaneously, we propose a Mamba-based Global Semantic Aggregator that injects semantics priors into keypoints to enhance their expressiveness and leverages multiple TwinMamba blocks to model long-range dependencies for more effective global feature aggregation. Extensive experiments on three benchmark datasets (REAL275, CAMERA25, and HouseCat6D) demonstrate that TSM-Pose outperforms existing state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16954
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TSM-Pose: Topology-Aware Learning with Semantic Mamba for Category-Level Object Pose Estimation
Liu, Jinshuo
Ma, Bingtao
Su, Junlin
Pan, Guanyuan
Wu, Beining
Yang, Cheng
Lu, Jiaxuan
Yan, Chenggang
Wang, Shuai
Computer Vision and Pattern Recognition
Category-level object pose estimation is fundamental for embodied intelligence, yet achieving robust generalization to unseen instances remains challenging. However, existing methods mainly rely on simple feature extraction and aggregation, which struggle to capture category-shared topological structures and conduct semantic keypoint modeling, limiting their generalization. To address these, we propose a \textbf{T}opology-Aware Learning with \textbf{S}emantic \textbf{M}amba for Category-Level \textbf{P}ose Estimation framework (TSM-Pose). Specifically, we introduce a Topology Extractor to capture the global topological representation of the point cloud, which is integrated into local geometry features and enables robust category-level structural representation. Simultaneously, we propose a Mamba-based Global Semantic Aggregator that injects semantics priors into keypoints to enhance their expressiveness and leverages multiple TwinMamba blocks to model long-range dependencies for more effective global feature aggregation. Extensive experiments on three benchmark datasets (REAL275, CAMERA25, and HouseCat6D) demonstrate that TSM-Pose outperforms existing state-of-the-art methods.
title TSM-Pose: Topology-Aware Learning with Semantic Mamba for Category-Level Object Pose Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.16954