Gespeichert in:
| Hauptverfasser: | Gan, Qijun, Ren, Yi, Zhang, Chen, Ye, Zhenhui, Xie, Pan, Yin, Xiang, Yuan, Zehuan, Peng, Bingyue, Zhu, Jianke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.04847 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
InfinityHuman: Towards Long-Term Audio-Driven Human
von: Li, Xiaodi, et al.
Veröffentlicht: (2025)
von: Li, Xiaodi, et al.
Veröffentlicht: (2025)
PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano Performance
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
HLLM: Enhancing Sequential Recommendations via Hierarchical Large Language Models for Item and User Modeling
von: Chen, Junyi, et al.
Veröffentlicht: (2024)
von: Chen, Junyi, et al.
Veröffentlicht: (2024)
ALIVE: Animate Your World with Lifelike Audio-Video Generation
von: Guo, Ying, et al.
Veröffentlicht: (2026)
von: Guo, Ying, et al.
Veröffentlicht: (2026)
Fine-Grained Multi-View Hand Reconstruction Using Inverse Rendering
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
XHand: Real-time Expressive Hand Avatar
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
HyperMotionX: The Dataset and Benchmark with DiT-Based Pose-Guided Human Image Animation of Complex Motions
von: Xu, Shuolin, et al.
Veröffentlicht: (2025)
von: Xu, Shuolin, et al.
Veröffentlicht: (2025)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
von: Tian, Keyu, et al.
Veröffentlicht: (2024)
von: Tian, Keyu, et al.
Veröffentlicht: (2024)
Generative Refinement Networks for Visual Synthesis
von: Han, Jian, et al.
Veröffentlicht: (2026)
von: Han, Jian, et al.
Veröffentlicht: (2026)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
von: Sun, Peize, et al.
Veröffentlicht: (2024)
von: Sun, Peize, et al.
Veröffentlicht: (2024)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
Do As I Do: Pose Guided Human Motion Copy
von: Wu, Sifan, et al.
Veröffentlicht: (2024)
von: Wu, Sifan, et al.
Veröffentlicht: (2024)
HLLM-Creator: Hierarchical LLM-based Personalized Creative Generation
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
VC-LLM: Automated Advertisement Video Creation from Raw Footage using Multi-modal LLMs
von: Qian, Dongjun, et al.
Veröffentlicht: (2025)
von: Qian, Dongjun, et al.
Veröffentlicht: (2025)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
von: Shao, Ruizhi, et al.
Veröffentlicht: (2024)
von: Shao, Ruizhi, et al.
Veröffentlicht: (2024)
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
von: Yang, Yuxiao, et al.
Veröffentlicht: (2025)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2025)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
von: Zhang, Shilong, et al.
Veröffentlicht: (2025)
von: Zhang, Shilong, et al.
Veröffentlicht: (2025)
Validation of Human Pose Estimation and Human Mesh Recovery for Extracting Clinically Relevant Motion Data from Videos
von: Armstrong, Kai, et al.
Veröffentlicht: (2025)
von: Armstrong, Kai, et al.
Veröffentlicht: (2025)
Language-Guided Transformer Tokenizer for Human Motion Generation
von: Yan, Sheng, et al.
Veröffentlicht: (2026)
von: Yan, Sheng, et al.
Veröffentlicht: (2026)
Waver: Wave Your Way to Lifelike Video Generation
von: Zhang, Yifu, et al.
Veröffentlicht: (2025)
von: Zhang, Yifu, et al.
Veröffentlicht: (2025)
HyperDiff: Hypergraph Guided Diffusion Model for 3D Human Pose Estimation
von: Han, Bing, et al.
Veröffentlicht: (2025)
von: Han, Bing, et al.
Veröffentlicht: (2025)
Rethinking Generative Human Video Coding with Implicit Motion Transformation
von: Chen, Bolin, et al.
Veröffentlicht: (2025)
von: Chen, Bolin, et al.
Veröffentlicht: (2025)
VITON-DiT: Learning In-the-Wild Video Try-On from Human Dance Videos via Diffusion Transformers
von: Zheng, Jun, et al.
Veröffentlicht: (2024)
von: Zheng, Jun, et al.
Veröffentlicht: (2024)
DiHuR: Diffusion-Guided Generalizable Human Reconstruction
von: Chen, Jinnan, et al.
Veröffentlicht: (2024)
von: Chen, Jinnan, et al.
Veröffentlicht: (2024)
Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
von: Qi, Tianhao, et al.
Veröffentlicht: (2025)
von: Qi, Tianhao, et al.
Veröffentlicht: (2025)
$\text{Di}^2\text{Pose}$: Discrete Diffusion Model for Occluded 3D Human Pose Estimation
von: Wang, Weiquan, et al.
Veröffentlicht: (2024)
von: Wang, Weiquan, et al.
Veröffentlicht: (2024)
AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models
von: Huang, Zehuan, et al.
Veröffentlicht: (2025)
von: Huang, Zehuan, et al.
Veröffentlicht: (2025)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
von: Ma, Chuofan, et al.
Veröffentlicht: (2025)
von: Ma, Chuofan, et al.
Veröffentlicht: (2025)
Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
von: Han, Jian, et al.
Veröffentlicht: (2024)
von: Han, Jian, et al.
Veröffentlicht: (2024)
HumanScore: Benchmarking Human Motions in Generated Videos
von: Fang, Yusu, et al.
Veröffentlicht: (2026)
von: Fang, Yusu, et al.
Veröffentlicht: (2026)
Diffusion-based Pose Refinement and Muti-hypothesis Generation for 3D Human Pose Estimaiton
von: Kang, Hongbo, et al.
Veröffentlicht: (2024)
von: Kang, Hongbo, et al.
Veröffentlicht: (2024)
Target Pose Guided Whole-body Grasping Motion Generation for Digital Humans
von: Shao, Quanquan, et al.
Veröffentlicht: (2024)
von: Shao, Quanquan, et al.
Veröffentlicht: (2024)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
von: He, Xuanhua, et al.
Veröffentlicht: (2025)
von: He, Xuanhua, et al.
Veröffentlicht: (2025)
Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
von: Chen, Pengtao, et al.
Veröffentlicht: (2025)
von: Chen, Pengtao, et al.
Veröffentlicht: (2025)
PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
von: He, Jingxuan, et al.
Veröffentlicht: (2025)
von: He, Jingxuan, et al.
Veröffentlicht: (2025)
SMooDi: Stylized Motion Diffusion Model
von: Zhong, Lei, et al.
Veröffentlicht: (2024)
von: Zhong, Lei, et al.
Veröffentlicht: (2024)
Kinematics Modeling Network for Video-based Human Pose Estimation
von: Dang, Yonghao, et al.
Veröffentlicht: (2022)
von: Dang, Yonghao, et al.
Veröffentlicht: (2022)
UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation
von: Jiang, Zhaodong, et al.
Veröffentlicht: (2025)
von: Jiang, Zhaodong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
InfinityHuman: Towards Long-Term Audio-Driven Human
von: Li, Xiaodi, et al.
Veröffentlicht: (2025) -
PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano Performance
von: Gan, Qijun, et al.
Veröffentlicht: (2024) -
HLLM: Enhancing Sequential Recommendations via Hierarchical Large Language Models for Item and User Modeling
von: Chen, Junyi, et al.
Veröffentlicht: (2024) -
ALIVE: Animate Your World with Lifelike Audio-Video Generation
von: Guo, Ying, et al.
Veröffentlicht: (2026) -
Fine-Grained Multi-View Hand Reconstruction Using Inverse Rendering
von: Gan, Qijun, et al.
Veröffentlicht: (2024)