UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xiang, Zhang, Shiwei, Tang, Longxiang, Zhang, Yingya, Gao, Changxin, Wang, Yuehuan, Sang, Nong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
von: Wang, Xiang, et al.
Veröffentlicht: (2024)
von: Wang, Xiang, et al.
Veröffentlicht: (2024)
Taming Consistency Distillation for Accelerated Human Image Animation
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
LLM2Manim: Pedagogy-Aware AI Generation of STEM Animations
von: Joshi, Aastha, et al.
Veröffentlicht: (2026)
von: Joshi, Aastha, et al.
Veröffentlicht: (2026)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Image Referenced Sketch Colorization Based on Animation Creation Workflow
von: Yan, Dingkun, et al.
Veröffentlicht: (2025)
von: Yan, Dingkun, et al.
Veröffentlicht: (2025)
AniME: Adaptive Multi-Agent Planning for Long Animation Generation
von: Zhang, Lisai, et al.
Veröffentlicht: (2025)
von: Zhang, Lisai, et al.
Veröffentlicht: (2025)
InstructHumans: Editing Animated 3D Human Textures with Instructions
von: Zhu, Jiayin, et al.
Veröffentlicht: (2024)
von: Zhu, Jiayin, et al.
Veröffentlicht: (2024)
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
Replace Anyone in Videos
von: Wang, Xiang, et al.
Veröffentlicht: (2024)
von: Wang, Xiang, et al.
Veröffentlicht: (2024)
Rethinking Vision Transformer for Large-Scale Fine-Grained Image Retrieval
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation
von: Lyu, Tianle, et al.
Veröffentlicht: (2025)
von: Lyu, Tianle, et al.
Veröffentlicht: (2025)
Every Painting Awakened: A Training-free Framework for Painting-to-Animation Generation
von: Liu, Lingyu, et al.
Veröffentlicht: (2025)
von: Liu, Lingyu, et al.
Veröffentlicht: (2025)
StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial Animation
von: Yang, An, et al.
Veröffentlicht: (2025)
von: Yang, An, et al.
Veröffentlicht: (2025)
RDTF: Resource-efficient Dual-mask Training Framework for Multi-frame Animated Sticker Generation
von: Yuan, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Yuan, Zhiqiang, et al.
Veröffentlicht: (2025)
Bridging the Gap: Sketch-Aware Interpolation Network for High-Quality Animation Sketch Inbetweening
von: Shen, Jiaming, et al.
Veröffentlicht: (2023)
von: Shen, Jiaming, et al.
Veröffentlicht: (2023)
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
von: Xu, Shuolin, et al.
Veröffentlicht: (2025)
von: Xu, Shuolin, et al.
Veröffentlicht: (2025)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
SAiD: Speech-driven Blendshape Facial Animation with Diffusion
von: Park, Inkyu, et al.
Veröffentlicht: (2023)
von: Park, Inkyu, et al.
Veröffentlicht: (2023)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
DanceCamAnimator: Keyframe-Based Controllable 3D Dance Camera Synthesis
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
2DGS-Avatar: Animatable High-fidelity Clothed Avatar via 2D Gaussian Splatting
von: Yan, Qipeng, et al.
Veröffentlicht: (2025)
von: Yan, Qipeng, et al.
Veröffentlicht: (2025)
KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation
von: Du, Guanyi, et al.
Veröffentlicht: (2026)
von: Du, Guanyi, et al.
Veröffentlicht: (2026)
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space
von: Liu, Kangwei, et al.
Veröffentlicht: (2025)
von: Liu, Kangwei, et al.
Veröffentlicht: (2025)
"You'll Be Alice Adventuring in Wonderland!" Processes, Challenges, and Opportunities of Creating Animated Virtual Reality Stories
von: Yuan, Lin-Ping, et al.
Veröffentlicht: (2025)
von: Yuan, Lin-Ping, et al.
Veröffentlicht: (2025)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
CLIP-guided Prototype Modulating for Few-shot Action Recognition
von: Wang, Xiang, et al.
Veröffentlicht: (2023)
von: Wang, Xiang, et al.
Veröffentlicht: (2023)
Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
von: Ji, Xiaozhong, et al.
Veröffentlicht: (2024)
von: Ji, Xiaozhong, et al.
Veröffentlicht: (2024)
Efficient and Accurate Image Provenance Analysis: A Scalable Pipeline for Large-scale Images
von: Lai, Jiewei, et al.
Veröffentlicht: (2025)
von: Lai, Jiewei, et al.
Veröffentlicht: (2025)
AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial Animation
von: Chen, Liyang, et al.
Veröffentlicht: (2023)
von: Chen, Liyang, et al.
Veröffentlicht: (2023)
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2026)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2026)
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding
von: Liu, Chang, et al.
Veröffentlicht: (2025)
von: Liu, Chang, et al.
Veröffentlicht: (2025)
Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
von: Zhou, Chao, et al.
Veröffentlicht: (2026)
von: Zhou, Chao, et al.
Veröffentlicht: (2026)
Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation
von: Huang, Zikai, et al.
Veröffentlicht: (2025)
von: Huang, Zikai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
von: Wang, Xiang, et al.
Veröffentlicht: (2024) -
Taming Consistency Distillation for Accelerated Human Image Animation
von: Wang, Xiang, et al.
Veröffentlicht: (2025) -
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
von: Qu, Qiang, et al.
Veröffentlicht: (2025) -
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
von: Sun, Jiahui, et al.
Veröffentlicht: (2025) -
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
von: Zhu, Rui, et al.
Veröffentlicht: (2024)