Waver: Wave Your Way to Lifelike Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yifu, Yang, Hao, Zhang, Yuqi, Hu, Yifei, Zhu, Fengda, Lin, Chuang, Mei, Xiaofeng, Jiang, Yi, Peng, Bingyue, Yuan, Zehuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ALIVE: Animate Your World with Lifelike Audio-Video Generation
by: Guo, Ying, et al.
Published: (2026)
by: Guo, Ying, et al.
Published: (2026)
InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
by: Liu, Jinlai, et al.
Published: (2025)
by: Liu, Jinlai, et al.
Published: (2025)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
Goku: Flow Based Video Generative Foundation Models
by: Chen, Shoufa, et al.
Published: (2025)
by: Chen, Shoufa, et al.
Published: (2025)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
by: Tian, Keyu, et al.
Published: (2024)
by: Tian, Keyu, et al.
Published: (2024)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
by: Sun, Peize, et al.
Published: (2024)
by: Sun, Peize, et al.
Published: (2024)
Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
by: Han, Jian, et al.
Published: (2024)
by: Han, Jian, et al.
Published: (2024)
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
HLLM: Enhancing Sequential Recommendations via Hierarchical Large Language Models for Item and User Modeling
by: Chen, Junyi, et al.
Published: (2024)
by: Chen, Junyi, et al.
Published: (2024)
Generative Refinement Networks for Visual Synthesis
by: Han, Jian, et al.
Published: (2026)
by: Han, Jian, et al.
Published: (2026)
Generative Region-Language Pretraining for Open-Ended Object Detection
by: Lin, Chuang, et al.
Published: (2024)
by: Lin, Chuang, et al.
Published: (2024)
HLLM-Creator: Hierarchical LLM-based Personalized Creative Generation
by: Chen, Junyi, et al.
Published: (2025)
by: Chen, Junyi, et al.
Published: (2025)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
by: Ma, Chuofan, et al.
Published: (2025)
by: Ma, Chuofan, et al.
Published: (2025)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
VC-LLM: Automated Advertisement Video Creation from Raw Footage using Multi-modal LLMs
by: Qian, Dongjun, et al.
Published: (2025)
by: Qian, Dongjun, et al.
Published: (2025)
Waverly Municipal Library: A User Survey.
by: Parr, Edward
Published: (1977)
by: Parr, Edward
Published: (1977)
InfinityHuman: Towards Long-Term Audio-Driven Human
by: Li, Xiaodi, et al.
Published: (2025)
by: Li, Xiaodi, et al.
Published: (2025)
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples
by: Hu, Zijing, et al.
Published: (2025)
by: Hu, Zijing, et al.
Published: (2025)
THGS: Lifelike Talking Human Avatar Synthesis From Monocular Video Via 3D Gaussian Splatting
by: Chuang Chen, et al.
Published: (2025)
by: Chuang Chen, et al.
Published: (2025)
FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
by: Bu, Jiazi, et al.
Published: (2024)
by: Bu, Jiazi, et al.
Published: (2024)
From Mannequin to Human: A Pose-Aware and Identity-Preserving Video Generation Framework for Lifelike Clothing Display
by: Mu, Xiangyu, et al.
Published: (2025)
by: Mu, Xiangyu, et al.
Published: (2025)
Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
by: Hu, Zijing, et al.
Published: (2025)
by: Hu, Zijing, et al.
Published: (2025)
Lifelike Agility and Play in Quadrupedal Robots using Reinforcement Learning and Generative Pre-trained Models
by: Han, Lei, et al.
Published: (2023)
by: Han, Lei, et al.
Published: (2023)
Veda: Scalable Video Diffusion via Distilled Sparse Attention
by: Han, Shihao, et al.
Published: (2026)
by: Han, Shihao, et al.
Published: (2026)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
by: Zhang, Jiaxu, et al.
Published: (2025)
by: Zhang, Jiaxu, et al.
Published: (2025)
UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control
by: Sun, Wenzhang, et al.
Published: (2024)
by: Sun, Wenzhang, et al.
Published: (2024)
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
by: Li, Ruibin, et al.
Published: (2025)
by: Li, Ruibin, et al.
Published: (2025)
Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
by: Li, Ruibin, et al.
Published: (2026)
by: Li, Ruibin, et al.
Published: (2026)
Fusion in Your Way: Aligning Image Fusion with Heterogeneous Demands via Direct Preference Optimization
by: Su, Weijian, et al.
Published: (2026)
by: Su, Weijian, et al.
Published: (2026)
VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
by: Xu, Sicheng, et al.
Published: (2024)
by: Xu, Sicheng, et al.
Published: (2024)
Motion Generation Review: Exploring Deep Learning for Lifelike Animation with Manifold
by: Zhao, Jiayi, et al.
Published: (2024)
by: Zhao, Jiayi, et al.
Published: (2024)
Towards Deviation-Robust Agent Navigation via Perturbation-Aware Contrastive Learning
by: Lin, Bingqian, et al.
Published: (2024)
by: Lin, Bingqian, et al.
Published: (2024)
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
by: Xu, Sicheng, et al.
Published: (2025)
by: Xu, Sicheng, et al.
Published: (2025)
Hierarchical‐Split Multi‐Scale Convolution Network With Multi‐Task Learning for Human Activity Recognition in Wearable Devices
by: Pengwei Zhang, et al.
Published: (2026)
by: Pengwei Zhang, et al.
Published: (2026)
Recognize Any Regions
by: Yang, Haosen, et al.
Published: (2023)
by: Yang, Haosen, et al.
Published: (2023)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
by: Ma, Chuofan, et al.
Published: (2024)
by: Ma, Chuofan, et al.
Published: (2024)
Finding Your Way with Your Baby
by: Daws, Dilys, et al.
Published: (2024)
by: Daws, Dilys, et al.
Published: (2024)
Pack and Force Your Memory: Long-form and Consistent Video Generation
by: Wu, Xiaofei, et al.
Published: (2025)
by: Wu, Xiaofei, et al.
Published: (2025)
Hyperion: Low-Latency Ultra-HD Video Analytics via Collaborative Vision Transformer Inference
by: Jiang, Linyi, et al.
Published: (2025)
by: Jiang, Linyi, et al.
Published: (2025)
Similar Items
-
ALIVE: Animate Your World with Lifelike Audio-Video Generation
by: Guo, Ying, et al.
Published: (2026) -
InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
by: Liu, Jinlai, et al.
Published: (2025) -
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
by: Zhang, Shilong, et al.
Published: (2025) -
Goku: Flow Based Video Generative Foundation Models
by: Chen, Shoufa, et al.
Published: (2025) -
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
by: Tian, Keyu, et al.
Published: (2024)