Generating Human Motion Videos using a Cascaded Text-to-Video Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Nam, Hyelin, Go, Hyojun, Park, Byeongjun, Kim, Byung-Hoon, Chung, Hyungjin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
by: Go, Hyojun, et al.
Published: (2025)
by: Go, Hyojun, et al.
Published: (2025)
SteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric Steering
by: Park, Byeongjun, et al.
Published: (2025)
by: Park, Byeongjun, et al.
Published: (2025)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
by: Chung, Hyungjin, et al.
Published: (2025)
by: Chung, Hyungjin, et al.
Published: (2025)
ReDirector: Creating Any-Length Video Retakes with Rotary Camera Encoding
by: Park, Byeongjun, et al.
Published: (2025)
by: Park, Byeongjun, et al.
Published: (2025)
Bridging Implicit and Explicit Geometric Transformation for Single-Image View Synthesis
by: Park, Byeongjun, et al.
Published: (2022)
by: Park, Byeongjun, et al.
Published: (2022)
Denoising Task Routing for Diffusion Models
by: Park, Byeongjun, et al.
Published: (2023)
by: Park, Byeongjun, et al.
Published: (2023)
Switch Diffusion Transformer: Synergizing Denoising Tasks with Sparse Mixture-of-Experts
by: Park, Byeongjun, et al.
Published: (2024)
by: Park, Byeongjun, et al.
Published: (2024)
SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis
by: Go, Hyojun, et al.
Published: (2024)
by: Go, Hyojun, et al.
Published: (2024)
CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language Models
by: Kim, Junho, et al.
Published: (2024)
by: Kim, Junho, et al.
Published: (2024)
CFG++: Manifold-constrained Classifier Free Guidance for Diffusion Models
by: Chung, Hyungjin, et al.
Published: (2024)
by: Chung, Hyungjin, et al.
Published: (2024)
Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
by: Go, Hyojun, et al.
Published: (2025)
by: Go, Hyojun, et al.
Published: (2025)
Diffusion Model Patching via Mixture-of-Prompts
by: Ham, Seokil, et al.
Published: (2024)
by: Ham, Seokil, et al.
Published: (2024)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
by: Taghipour, Ashkan, et al.
Published: (2026)
by: Taghipour, Ashkan, et al.
Published: (2026)
Optical-Flow Guided Prompt Optimization for Coherent Video Generation
by: Nam, Hyelin, et al.
Published: (2024)
by: Nam, Hyelin, et al.
Published: (2024)
Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift
by: Kim, Gihoon, et al.
Published: (2025)
by: Kim, Gihoon, et al.
Published: (2025)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
by: Qian, Zhaofang, et al.
Published: (2024)
by: Qian, Zhaofang, et al.
Published: (2024)
Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models
by: Kwon, Taesung, et al.
Published: (2026)
by: Kwon, Taesung, et al.
Published: (2026)
MotionBooth: Motion-Aware Customized Text-to-Video Generation
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
Contrastive Denoising Score for Text-guided Latent Diffusion Image Editing
by: Nam, Hyelin, et al.
Published: (2023)
by: Nam, Hyelin, et al.
Published: (2023)
Fleximo: Towards Flexible Text-to-Human Motion Video Generation
by: Zhang, Yuhang, et al.
Published: (2024)
by: Zhang, Yuhang, et al.
Published: (2024)
HumanScore: Benchmarking Human Motions in Generated Videos
by: Fang, Yusu, et al.
Published: (2026)
by: Fang, Yusu, et al.
Published: (2026)
MEVG: Multi-event Video Generation with Text-to-Video Models
by: Oh, Gyeongrok, et al.
Published: (2023)
by: Oh, Gyeongrok, et al.
Published: (2023)
Motion by Queries: Identity-Motion Trade-offs in Text-to-Video Generation
by: Atzmon, Yuval, et al.
Published: (2024)
by: Atzmon, Yuval, et al.
Published: (2024)
ContextMRI: Enhancing Compressed Sensing MRI through Metadata Conditioning
by: Chung, Hyungjin, et al.
Published: (2025)
by: Chung, Hyungjin, et al.
Published: (2025)
One Look is Enough: Seamless Patchwise Refinement for Zero-Shot Monocular Depth Estimation on High-Resolution Images
by: Kwon, Byeongjun, et al.
Published: (2025)
by: Kwon, Byeongjun, et al.
Published: (2025)
MATRIX: Mask Track Alignment for Interaction-aware Video Generation
by: Jin, Siyoon, et al.
Published: (2025)
by: Jin, Siyoon, et al.
Published: (2025)
MotionCharacter: Fine-Grained Motion Controllable Human Video Generation
by: Fang, Haopeng, et al.
Published: (2024)
by: Fang, Haopeng, et al.
Published: (2024)
MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
by: Yang, Kaixing, et al.
Published: (2025)
by: Yang, Kaixing, et al.
Published: (2025)
CoMo: Compositional Motion Customization for Text-to-Video Generation
by: Xu, Youcan, et al.
Published: (2025)
by: Xu, Youcan, et al.
Published: (2025)
Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and Conditioning
by: Ruan, Penghui, et al.
Published: (2024)
by: Ruan, Penghui, et al.
Published: (2024)
Regularization by Texts for Latent Diffusion Inverse Solvers
by: Kim, Jeongsol, et al.
Published: (2023)
by: Kim, Jeongsol, et al.
Published: (2023)
Human Motion Video Generation: A Survey
by: Xue, Haiwei, et al.
Published: (2025)
by: Xue, Haiwei, et al.
Published: (2025)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025)
by: Jeong, Boseung, et al.
Published: (2025)
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
by: Wang, Boyuan, et al.
Published: (2025)
by: Wang, Boyuan, et al.
Published: (2025)
FlashVideo: A Framework for Swift Inference in Text-to-Video Generation
by: Lei, Bin, et al.
Published: (2023)
by: Lei, Bin, et al.
Published: (2023)
MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
by: Liu, Yanchen, et al.
Published: (2025)
by: Liu, Yanchen, et al.
Published: (2025)
Toward Rich Video Human-Motion2D Generation
by: Xi, Ruihao, et al.
Published: (2025)
by: Xi, Ruihao, et al.
Published: (2025)
Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments
by: Kwon, Soonwoo, et al.
Published: (2025)
by: Kwon, Soonwoo, et al.
Published: (2025)
Text-based Talking Video Editing with Cascaded Conditional Diffusion
by: Han, Bo, et al.
Published: (2024)
by: Han, Bo, et al.
Published: (2024)
Similar Items
-
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
by: Go, Hyojun, et al.
Published: (2025) -
SteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric Steering
by: Park, Byeongjun, et al.
Published: (2025) -
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
by: Chung, Hyungjin, et al.
Published: (2025) -
ReDirector: Creating Any-Length Video Retakes with Rotary Camera Encoding
by: Park, Byeongjun, et al.
Published: (2025) -
Bridging Implicit and Explicit Geometric Transformation for Single-Image View Synthesis
by: Park, Byeongjun, et al.
Published: (2022)