CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918248723775488 |
|---|---|
| author | Zhu, Yu Zeng, Dan Li, Shuiwang Zhao, Qijun Shen, Qiaomu Tang, Bo |
| author_facet | Zhu, Yu Zeng, Dan Li, Shuiwang Zhao, Qijun Shen, Qiaomu Tang, Bo |
| contents | Recent research in Category-Agnostic Pose Estimation (CAPE) has adopted fixed textual keypoint description as semantic prior for two-stage pose matching frameworks. While this paradigm enhances robustness and flexibility by disentangling the dependency of support images, our critical analysis reveals two inherent limitations of static joint embedding: (1) polysemy-induced cross-category ambiguity during the matching process(e.g., the concept "leg" exhibiting divergent visual manifestations across humans and furniture), and (2) insufficient discriminability for fine-grained intra-category variations (e.g., posture and fur discrepancies between a sleeping white cat and a standing black cat). To overcome these challenges, we propose a new framework that innovatively integrates hierarchical cross-modal interaction with dual-stream feature refinement, enhancing the joint embedding with both class-level and instance-specific cues from textual description and specific images. Experiments on the MP-100 dataset demonstrate that, regardless of the network backbone, CapeNext consistently outperforms state-of-the-art CAPE methods by a large margin. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_13102 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation Zhu, Yu Zeng, Dan Li, Shuiwang Zhao, Qijun Shen, Qiaomu Tang, Bo Computer Vision and Pattern Recognition Recent research in Category-Agnostic Pose Estimation (CAPE) has adopted fixed textual keypoint description as semantic prior for two-stage pose matching frameworks. While this paradigm enhances robustness and flexibility by disentangling the dependency of support images, our critical analysis reveals two inherent limitations of static joint embedding: (1) polysemy-induced cross-category ambiguity during the matching process(e.g., the concept "leg" exhibiting divergent visual manifestations across humans and furniture), and (2) insufficient discriminability for fine-grained intra-category variations (e.g., posture and fur discrepancies between a sleeping white cat and a standing black cat). To overcome these challenges, we propose a new framework that innovatively integrates hierarchical cross-modal interaction with dual-stream feature refinement, enhancing the joint embedding with both class-level and instance-specific cues from textual description and specific images. Experiments on the MP-100 dataset demonstrate that, regardless of the network backbone, CapeNext consistently outperforms state-of-the-art CAPE methods by a large margin. |
| title | CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.13102 |