Rethink Sparse Signals for Pose-guided Text-to-image Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xuan, Wenjie, Zhang, Jing, Liu, Juhua, Du, Bo, Tao, Dacheng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915359644188672
author Xuan, Wenjie
Zhang, Jing
Liu, Juhua
Du, Bo
Tao, Dacheng
author_facet Xuan, Wenjie
Zhang, Jing
Liu, Juhua
Du, Bo
Tao, Dacheng
contents Recent works favored dense signals (e.g., depth, DensePose), as an alternative to sparse signals (e.g., OpenPose), to provide detailed spatial guidance for pose-guided text-to-image generation. However, dense representations raised new challenges, including editing difficulties and potential inconsistencies with textual prompts. This fact motivates us to revisit sparse signals for pose guidance, owing to their simplicity and shape-agnostic nature, which remains underexplored. This paper proposes a novel Spatial-Pose ControlNet(SP-Ctrl), equipping sparse signals with robust controllability for pose-guided image generation. Specifically, we extend OpenPose to a learnable spatial representation, making keypoint embeddings discriminative and expressive. Additionally, we introduce keypoint concept learning, which encourages keypoint tokens to attend to the spatial positions of each keypoint, thus improving pose alignment. Experiments on animal- and human-centric image generation tasks demonstrate that our method outperforms recent spatially controllable T2I generation approaches under sparse-pose guidance and even matches the performance of dense signal-based methods. Moreover, SP-Ctrl shows promising capabilities in diverse and cross-species generation through sparse signals. Codes will be available at https://github.com/DREAMXFAR/SP-Ctrl.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20983
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethink Sparse Signals for Pose-guided Text-to-image Generation
Xuan, Wenjie
Zhang, Jing
Liu, Juhua
Du, Bo
Tao, Dacheng
Computer Vision and Pattern Recognition
Recent works favored dense signals (e.g., depth, DensePose), as an alternative to sparse signals (e.g., OpenPose), to provide detailed spatial guidance for pose-guided text-to-image generation. However, dense representations raised new challenges, including editing difficulties and potential inconsistencies with textual prompts. This fact motivates us to revisit sparse signals for pose guidance, owing to their simplicity and shape-agnostic nature, which remains underexplored. This paper proposes a novel Spatial-Pose ControlNet(SP-Ctrl), equipping sparse signals with robust controllability for pose-guided image generation. Specifically, we extend OpenPose to a learnable spatial representation, making keypoint embeddings discriminative and expressive. Additionally, we introduce keypoint concept learning, which encourages keypoint tokens to attend to the spatial positions of each keypoint, thus improving pose alignment. Experiments on animal- and human-centric image generation tasks demonstrate that our method outperforms recent spatially controllable T2I generation approaches under sparse-pose guidance and even matches the performance of dense signal-based methods. Moreover, SP-Ctrl shows promising capabilities in diverse and cross-species generation through sparse signals. Codes will be available at https://github.com/DREAMXFAR/SP-Ctrl.
title Rethink Sparse Signals for Pose-guided Text-to-image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.20983