ChildlikeSHAPES: Semantic Hierarchical Region Parsing for Animating Figure Drawings

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Srivastava, Astitva, Smith, Harrison Jesse, Nguyen-Phuoc, Thu, Ye, Yuting
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908313142165504
author Srivastava, Astitva
Smith, Harrison Jesse
Nguyen-Phuoc, Thu
Ye, Yuting
author_facet Srivastava, Astitva
Smith, Harrison Jesse
Nguyen-Phuoc, Thu
Ye, Yuting
contents Childlike human figure drawings represent one of humanity's most accessible forms of character expression, yet automatically analyzing their contents remains a significant challenge. While semantic segmentation of realistic humans has recently advanced considerably, existing models often fail when confronted with the abstract, representational nature of childlike drawings. This semantic understanding is a crucial prerequisite for animation tools that seek to modify figures while preserving their unique style. To help achieve this, we propose a novel hierarchical segmentation model, built upon the architecture and pre-trained SAM, to quickly and accurately obtain these semantic labels. Our model achieves higher accuracy than state-of-the-art segmentation models focused on realistic humans and cartoon figures, even after fine-tuning. We demonstrate the value of our model for semantic segmentation through multiple applications: a fully automatic facial animation pipeline, a figure relighting pipeline, improvements to an existing childlike human figure drawing animation method, and generalization to out-of-domain figures. Finally, to support future work in this area, we introduce a dataset of 16,000 childlike drawings with pixel-level annotations across 25 semantic categories. Our work can enable entirely new, easily accessible tools for hand-drawn character animation, and our dataset can enable new lines of inquiry in a variety of graphics and human-centric research fields.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08022
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChildlikeSHAPES: Semantic Hierarchical Region Parsing for Animating Figure Drawings
Srivastava, Astitva
Smith, Harrison Jesse
Nguyen-Phuoc, Thu
Ye, Yuting
Graphics
Childlike human figure drawings represent one of humanity's most accessible forms of character expression, yet automatically analyzing their contents remains a significant challenge. While semantic segmentation of realistic humans has recently advanced considerably, existing models often fail when confronted with the abstract, representational nature of childlike drawings. This semantic understanding is a crucial prerequisite for animation tools that seek to modify figures while preserving their unique style. To help achieve this, we propose a novel hierarchical segmentation model, built upon the architecture and pre-trained SAM, to quickly and accurately obtain these semantic labels. Our model achieves higher accuracy than state-of-the-art segmentation models focused on realistic humans and cartoon figures, even after fine-tuning. We demonstrate the value of our model for semantic segmentation through multiple applications: a fully automatic facial animation pipeline, a figure relighting pipeline, improvements to an existing childlike human figure drawing animation method, and generalization to out-of-domain figures. Finally, to support future work in this area, we introduce a dataset of 16,000 childlike drawings with pixel-level annotations across 25 semantic categories. Our work can enable entirely new, easily accessible tools for hand-drawn character animation, and our dataset can enable new lines of inquiry in a variety of graphics and human-centric research fields.
title ChildlikeSHAPES: Semantic Hierarchical Region Parsing for Animating Figure Drawings
topic Graphics
url https://arxiv.org/abs/2504.08022