OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hui, Xu, Mingwang, Zhan, Yun, Mu, Shan, Li, Jiaye, Cheng, Kaihui, Chen, Yuxuan, Chen, Tan, Ye, Mao, Wang, Jingdong, Zhu, Siyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
by: Li, Chunyu, et al.
Published: (2026)
by: Li, Chunyu, et al.
Published: (2026)
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
by: Zhu, Lei, et al.
Published: (2026)
by: Zhu, Lei, et al.
Published: (2026)
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
by: Li, Jiaye, et al.
Published: (2025)
by: Li, Jiaye, et al.
Published: (2025)
Pyramidal Patchification Flow for Visual Generation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
by: Nan, Kepan, et al.
Published: (2024)
by: Nan, Kepan, et al.
Published: (2024)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025)
by: Qiu, Zongyang, et al.
Published: (2025)
MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers
by: Yao, Yuxuan, et al.
Published: (2026)
by: Yao, Yuxuan, et al.
Published: (2026)
Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization
by: Cui, Jiahao, et al.
Published: (2025)
by: Cui, Jiahao, et al.
Published: (2025)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation
by: Zhang, Youliang, et al.
Published: (2025)
by: Zhang, Youliang, et al.
Published: (2025)
DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
by: Di, Donglin, et al.
Published: (2024)
by: Di, Donglin, et al.
Published: (2024)
Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive Prompting
by: Liu, Yuanyuan, et al.
Published: (2024)
by: Liu, Yuanyuan, et al.
Published: (2024)
HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation
by: Wang, Zhenzhi, et al.
Published: (2024)
by: Wang, Zhenzhi, et al.
Published: (2024)
VidAnimator: User-Guided Stylized 3D Character Animation from Human Videos
by: Ye, Xinwu, et al.
Published: (2025)
by: Ye, Xinwu, et al.
Published: (2025)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
by: Cai, Yuxuan, et al.
Published: (2025)
by: Cai, Yuxuan, et al.
Published: (2025)
VidGen-1M: A Large-Scale Dataset for Text-to-video Generation
by: Tan, Zhiyu, et al.
Published: (2024)
by: Tan, Zhiyu, et al.
Published: (2024)
UniVid: The Open-Source Unified Video Model
by: Luo, Jiabin, et al.
Published: (2025)
by: Luo, Jiabin, et al.
Published: (2025)
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
by: Zhou, Ting, et al.
Published: (2024)
by: Zhou, Ting, et al.
Published: (2024)
TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
by: Chen, Shunian, et al.
Published: (2025)
by: Chen, Shunian, et al.
Published: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
VidTok: A Versatile and Open-Source Video Tokenizer
by: Tang, Anni, et al.
Published: (2024)
by: Tang, Anni, et al.
Published: (2024)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
by: Chen, Lan, et al.
Published: (2025)
by: Chen, Lan, et al.
Published: (2025)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
by: Xiao, Xinyu, et al.
Published: (2026)
by: Xiao, Xinyu, et al.
Published: (2026)
Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image Generation
by: Wang, Junyan, et al.
Published: (2024)
by: Wang, Junyan, et al.
Published: (2024)
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
by: Li, Keliang, et al.
Published: (2025)
by: Li, Keliang, et al.
Published: (2025)
VidLA: Video-Language Alignment at Scale
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
World In Your Hands: A Large-Scale and Open-Source Ecosystem for Learning Human-Centric Manipulation in the Wild
by: Zheng, Yupeng, et al.
Published: (2025)
by: Zheng, Yupeng, et al.
Published: (2025)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
by: Zhang, Zhicheng, et al.
Published: (2025)
by: Zhang, Zhicheng, et al.
Published: (2025)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
by: Li, Keliang, et al.
Published: (2024)
by: Li, Keliang, et al.
Published: (2024)
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
by: Xu, Mingwang, et al.
Published: (2024)
by: Xu, Mingwang, et al.
Published: (2024)
VidLeaks: Membership Inference Attacks Against Text-to-Video Models
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
LLM-Guided Multi-View Hypergraph Learning for Human-Centric Explainable Recommendation
by: Chu, Zhixuan, et al.
Published: (2024)
by: Chu, Zhixuan, et al.
Published: (2024)
A Cross‐Organ Symphony of Human Immune System Maturation
by: Zhan‐Li Chen
Published: (2026)
by: Zhan‐Li Chen
Published: (2026)
HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions
by: Dong, Yifei, et al.
Published: (2025)
by: Dong, Yifei, et al.
Published: (2025)
Similar Items
-
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
by: Li, Chunyu, et al.
Published: (2026) -
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
by: Cui, Jiahao, et al.
Published: (2024) -
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
by: Zhu, Lei, et al.
Published: (2026) -
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
by: Li, Jiaye, et al.
Published: (2025) -
Pyramidal Patchification Flow for Visual Generation
by: Li, Hui, et al.
Published: (2025)