A Comprehensive Survey on Human Video Generation: Challenges, Methods, and Insights
Fuente:
arXiv
Saved in:
| Main Authors: | Lei, Wentao, Wang, Jinting, Ma, Fengji, Huang, Guanjie, Liu, Li |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability
by: Ma, Fengji, et al.
Published: (2024)
by: Ma, Fengji, et al.
Published: (2024)
Segment Anything for Videos: A Systematic Survey
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
A Comprehensive Survey on 3D Content Generation
by: Liu, Jian, et al.
Published: (2024)
by: Liu, Jian, et al.
Published: (2024)
Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges
by: Akter, Sanjeda, et al.
Published: (2025)
by: Akter, Sanjeda, et al.
Published: (2025)
Diffusion Models: A Comprehensive Survey of Methods and Applications
by: Yang, Ling, et al.
Published: (2022)
by: Yang, Ling, et al.
Published: (2022)
A Survey: Spatiotemporal Consistency in Video Generation
by: Yin, Zhiyu, et al.
Published: (2025)
by: Yin, Zhiyu, et al.
Published: (2025)
Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
by: Li, Jinxuan, et al.
Published: (2025)
by: Li, Jinxuan, et al.
Published: (2025)
Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body
by: Wang, Zeqing, et al.
Published: (2024)
by: Wang, Zeqing, et al.
Published: (2024)
InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative Refinement
by: Zou, Yude, et al.
Published: (2026)
by: Zou, Yude, et al.
Published: (2026)
SAM2 for Image and Video Segmentation: A Comprehensive Survey
by: Jiaxing, Zhang, et al.
Published: (2025)
by: Jiaxing, Zhang, et al.
Published: (2025)
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
by: Wei, Yuancheng, et al.
Published: (2026)
by: Wei, Yuancheng, et al.
Published: (2026)
HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks
by: Zhang, Fengji, et al.
Published: (2024)
by: Zhang, Fengji, et al.
Published: (2024)
A Comprehensive Survey of Data Augmentation in Visual Reinforcement Learning
by: Ma, Guozheng, et al.
Published: (2022)
by: Ma, Guozheng, et al.
Published: (2022)
WebUOT-1M: Advancing Deep Underwater Object Tracking with A Million-Scale Benchmark
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Video-Bench: Human-Aligned Video Generation Benchmark
by: Han, Hui, et al.
Published: (2025)
by: Han, Hui, et al.
Published: (2025)
Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation
by: Lei, Yinjie, et al.
Published: (2023)
by: Lei, Yinjie, et al.
Published: (2023)
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions
by: Dong, Yifei, et al.
Published: (2025)
by: Dong, Yifei, et al.
Published: (2025)
Challenges and Trends in Egocentric Vision: A Survey
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
by: Zhou, Pengyuan, et al.
Published: (2024)
by: Zhou, Pengyuan, et al.
Published: (2024)
CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
by: Feng, Huidong, et al.
Published: (2026)
by: Feng, Huidong, et al.
Published: (2026)
A Survey of Adversarial Defenses in Vision-based Systems: Categorization, Methods and Challenges
by: Chattopadhyay, Nandish, et al.
Published: (2025)
by: Chattopadhyay, Nandish, et al.
Published: (2025)
Exploring the Role of Synthetic Data Augmentation in Controllable Human-Centric Video Generation
by: Fei, Yuanchen, et al.
Published: (2026)
by: Fei, Yuanchen, et al.
Published: (2026)
EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation
by: Zhang, Libo, et al.
Published: (2025)
by: Zhang, Libo, et al.
Published: (2025)
Distorted or Fabricated? A Survey on Hallucination in Video LLMs
by: Huang, Yiyang, et al.
Published: (2026)
by: Huang, Yiyang, et al.
Published: (2026)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
by: Zhang, Zhihong, et al.
Published: (2025)
by: Zhang, Zhihong, et al.
Published: (2025)
A Comprehensive Survey of Continual Learning: Theory, Method and Application
by: Wang, Liyuan, et al.
Published: (2023)
by: Wang, Liyuan, et al.
Published: (2023)
AuthSig: Safeguarding Scanned Signatures Against Unauthorized Reuse in Paperless Workflows
by: Zhang, RuiQiang, et al.
Published: (2025)
by: Zhang, RuiQiang, et al.
Published: (2025)
Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D Generation
by: Li, Ziying, et al.
Published: (2025)
by: Li, Ziying, et al.
Published: (2025)
Draft-and-Target Sampling for Video Generation Policy
by: Zhang, Qikang, et al.
Published: (2026)
by: Zhang, Qikang, et al.
Published: (2026)
Towards Fine-Grained Human Motion Video Captioning
by: Song, Guorui, et al.
Published: (2025)
by: Song, Guorui, et al.
Published: (2025)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
by: Li, Changlin, et al.
Published: (2025)
by: Li, Changlin, et al.
Published: (2025)
Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?
by: Li, Hongyu, et al.
Published: (2026)
by: Li, Hongyu, et al.
Published: (2026)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
by: Luo, Yongdong, et al.
Published: (2024)
by: Luo, Yongdong, et al.
Published: (2024)
Vision Mamba: A Comprehensive Survey and Taxonomy
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
A Survey on Long Video Generation: Challenges, Methods, and Prospects
by: Li, Chengxuan, et al.
Published: (2024)
by: Li, Chengxuan, et al.
Published: (2024)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
ATI: Any Trajectory Instruction for Controllable Video Generation
by: Wang, Angtian, et al.
Published: (2025)
by: Wang, Angtian, et al.
Published: (2025)
Similar Items
-
TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability
by: Ma, Fengji, et al.
Published: (2024) -
Segment Anything for Videos: A Systematic Survey
by: Zhang, Chunhui, et al.
Published: (2024) -
A Comprehensive Survey on 3D Content Generation
by: Liu, Jian, et al.
Published: (2024) -
Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges
by: Akter, Sanjeda, et al.
Published: (2025) -
Diffusion Models: A Comprehensive Survey of Methods and Applications
by: Yang, Ling, et al.
Published: (2022)