SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yan, Wenhao, Ye, Sheng, Yang, Zhuoyi, Teng, Jiayan, Dong, ZhenHui, Wen, Kairui, Gu, Xiaotao, Liu, Yong-Jin, Tang, Jie
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912978235817984
author Yan, Wenhao
Ye, Sheng
Yang, Zhuoyi
Teng, Jiayan
Dong, ZhenHui
Wen, Kairui
Gu, Xiaotao
Liu, Yong-Jin
Tang, Jie
author_facet Yan, Wenhao
Ye, Sheng
Yang, Zhuoyi
Teng, Jiayan
Dong, ZhenHui
Wen, Kairui
Gu, Xiaotao
Liu, Yong-Jin
Tang, Jie
contents Achieving controllable character animation that meets studio-grade standards remains challenging despite recent progress. Existing approaches can transfer motion from a driving video to a reference image, but often fail to preserve structural fidelity and temporal consistency in wild scenarios involving complex motion and cross-identity animations. In this work, we present \textbf{SCAIL} (a framework toward \textbf{S}tudio-grade \textbf{C}haracter \textbf{A}nimation via \textbf{I}n-context \textbf{L}earning), which is designed to address these challenges from two key innovations. First, we propose a novel 3D pose representation, providing a robust and flexible motion signal. Second, we introduce a full-context pose injection mechanism within a diffusion-transformer, enabling effective spatio-temporal reasoning over full motion sequences. To align with studio-grade requirements, we develop a curated data pipeline ensuring both diversity and quality, and establish a comprehensive benchmark for systematic evaluation. Experiments show that \textbf{SCAIL} achieves state-of-the-art performance and advances character animation toward studio-grade controlling. Code and model are available at \href{https://github.com/zai-org/SCAIL}{zai-org/SCAIL}.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05905
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
Yan, Wenhao
Ye, Sheng
Yang, Zhuoyi
Teng, Jiayan
Dong, ZhenHui
Wen, Kairui
Gu, Xiaotao
Liu, Yong-Jin
Tang, Jie
Computer Vision and Pattern Recognition
Achieving controllable character animation that meets studio-grade standards remains challenging despite recent progress. Existing approaches can transfer motion from a driving video to a reference image, but often fail to preserve structural fidelity and temporal consistency in wild scenarios involving complex motion and cross-identity animations. In this work, we present \textbf{SCAIL} (a framework toward \textbf{S}tudio-grade \textbf{C}haracter \textbf{A}nimation via \textbf{I}n-context \textbf{L}earning), which is designed to address these challenges from two key innovations. First, we propose a novel 3D pose representation, providing a robust and flexible motion signal. Second, we introduce a full-context pose injection mechanism within a diffusion-transformer, enabling effective spatio-temporal reasoning over full motion sequences. To align with studio-grade requirements, we develop a curated data pipeline ensuring both diversity and quality, and establish a comprehensive benchmark for systematic evaluation. Experiments show that \textbf{SCAIL} achieves state-of-the-art performance and advances character animation toward studio-grade controlling. Code and model are available at \href{https://github.com/zai-org/SCAIL}{zai-org/SCAIL}.
title SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.05905