MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators
Fuente:
arXiv
Salvato in:
| Autori principali: | Yuan, Shenghai, Huang, Jinfa, Shi, Yujun, Xu, Yongqi, Zhu, Ruijie, Lin, Bin, Cheng, Xinhua, Yuan, Li, Luo, Jiebo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
di: Yuan, Shenghai, et al.
Pubblicazione: (2024)
di: Yuan, Shenghai, et al.
Pubblicazione: (2024)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
di: Yuan, Shenghai, et al.
Pubblicazione: (2024)
di: Yuan, Shenghai, et al.
Pubblicazione: (2024)
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
di: Yuan, Shenghai, et al.
Pubblicazione: (2025)
di: Yuan, Shenghai, et al.
Pubblicazione: (2025)
FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation
di: Ge, Yunyang, et al.
Pubblicazione: (2025)
di: Ge, Yunyang, et al.
Pubblicazione: (2025)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
di: Li, Zongjian, et al.
Pubblicazione: (2024)
di: Li, Zongjian, et al.
Pubblicazione: (2024)
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
di: Chen, Liuhan, et al.
Pubblicazione: (2024)
di: Chen, Liuhan, et al.
Pubblicazione: (2024)
Helios: Real Real-Time Long Video Generation Model
di: Yuan, Shenghai, et al.
Pubblicazione: (2026)
di: Yuan, Shenghai, et al.
Pubblicazione: (2026)
HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
di: Zhou, Haiyang, et al.
Pubblicazione: (2025)
di: Zhou, Haiyang, et al.
Pubblicazione: (2025)
OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning
di: Ge, Yunyang, et al.
Pubblicazione: (2026)
di: Ge, Yunyang, et al.
Pubblicazione: (2026)
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
di: Wang, Yongqi, et al.
Pubblicazione: (2024)
di: Wang, Yongqi, et al.
Pubblicazione: (2024)
Evolver: Chain-of-Evolution Prompting to Boost Large Multimodal Models for Hateful Meme Detection
di: Huang, Jinfa, et al.
Pubblicazione: (2024)
di: Huang, Jinfa, et al.
Pubblicazione: (2024)
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
di: Wang, Xiaodong, et al.
Pubblicazione: (2025)
di: Wang, Xiaodong, et al.
Pubblicazione: (2025)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
MonarchRT: Efficient Attention for Real-Time Video Generation
di: Agarwal, Krish, et al.
Pubblicazione: (2026)
di: Agarwal, Krish, et al.
Pubblicazione: (2026)
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
di: Huang, Haoyu, et al.
Pubblicazione: (2026)
di: Huang, Haoyu, et al.
Pubblicazione: (2026)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
di: Luo, Yongdong, et al.
Pubblicazione: (2024)
di: Luo, Yongdong, et al.
Pubblicazione: (2024)
QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension
di: Luo, Yongdong, et al.
Pubblicazione: (2025)
di: Luo, Yongdong, et al.
Pubblicazione: (2025)
SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video
di: Zhao, Chengshu, et al.
Pubblicazione: (2025)
di: Zhao, Chengshu, et al.
Pubblicazione: (2025)
GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning
di: Li, Yujun, et al.
Pubblicazione: (2026)
di: Li, Yujun, et al.
Pubblicazione: (2026)
Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based Approach
di: Zhang, Shaofeng, et al.
Pubblicazione: (2024)
di: Zhang, Shaofeng, et al.
Pubblicazione: (2024)
DreamStereo: Towards Real-Time Stereo Inpainting for HD Videos
di: Huang, Yuan, et al.
Pubblicazione: (2026)
di: Huang, Yuan, et al.
Pubblicazione: (2026)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
di: Lin, Bin, et al.
Pubblicazione: (2025)
di: Lin, Bin, et al.
Pubblicazione: (2025)
DiffMagicFace: Identity Consistent Facial Editing of Real Videos
di: Yin, Huanghao, et al.
Pubblicazione: (2026)
di: Yin, Huanghao, et al.
Pubblicazione: (2026)
Open-Sora Plan: Open-Source Large Video Generation Model
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
di: Zhou, Zhenghong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenghong, et al.
Pubblicazione: (2024)
Jacquard V2: Refining Datasets using the Human In the Loop Data Correction Method
di: Li, Qiuhao, et al.
Pubblicazione: (2024)
di: Li, Qiuhao, et al.
Pubblicazione: (2024)
Interactive Test-Time Adaptation with Reliable Spatial-Temporal Voxels for Multi-Modal Segmentation
di: Cao, Haozhi, et al.
Pubblicazione: (2024)
di: Cao, Haozhi, et al.
Pubblicazione: (2024)
MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers
di: Zhang, Yuechen, et al.
Pubblicazione: (2025)
di: Zhang, Yuechen, et al.
Pubblicazione: (2025)
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives
di: Meng, Yihao, et al.
Pubblicazione: (2026)
di: Meng, Yihao, et al.
Pubblicazione: (2026)
DATE: Dynamic Absolute Time Enhancement for Long Video Understanding
di: Yuan, Chao, et al.
Pubblicazione: (2025)
di: Yuan, Chao, et al.
Pubblicazione: (2025)
MagicFight: Personalized Martial Arts Combat Video Generation
di: Huang, Jiancheng, et al.
Pubblicazione: (2026)
di: Huang, Jiancheng, et al.
Pubblicazione: (2026)
Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
di: Tang, Zhenyu, et al.
Pubblicazione: (2024)
di: Tang, Zhenyu, et al.
Pubblicazione: (2024)
Real-Time Motion-Controllable Autoregressive Video Diffusion
di: Zhao, Kesen, et al.
Pubblicazione: (2025)
di: Zhao, Kesen, et al.
Pubblicazione: (2025)
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
di: Li, Zongjian, et al.
Pubblicazione: (2025)
di: Li, Zongjian, et al.
Pubblicazione: (2025)
Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation
di: Chen, Wenting, et al.
Pubblicazione: (2023)
di: Chen, Wenting, et al.
Pubblicazione: (2023)
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
di: Yao, Yuan, et al.
Pubblicazione: (2025)
di: Yao, Yuan, et al.
Pubblicazione: (2025)
Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos
di: Alzayer, Hadi, et al.
Pubblicazione: (2024)
di: Alzayer, Hadi, et al.
Pubblicazione: (2024)
Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
Magic 1-For-1: Generating One Minute Video Clips within One Minute
di: Yi, Hongwei, et al.
Pubblicazione: (2025)
di: Yi, Hongwei, et al.
Pubblicazione: (2025)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
di: Hua, Hang, et al.
Pubblicazione: (2024)
di: Hua, Hang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
di: Yuan, Shenghai, et al.
Pubblicazione: (2024) -
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
di: Yuan, Shenghai, et al.
Pubblicazione: (2024) -
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
di: Yuan, Shenghai, et al.
Pubblicazione: (2025) -
FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation
di: Ge, Yunyang, et al.
Pubblicazione: (2025) -
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
di: Li, Zongjian, et al.
Pubblicazione: (2024)