VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yumeng, Beluch, William, Keuper, Margret, Zhang, Dan, Khoreva, Anna |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Divide & Bind Your Attention for Improved Generative Semantic Nursing
di: Li, Yumeng, et al.
Pubblicazione: (2023)
di: Li, Yumeng, et al.
Pubblicazione: (2023)
Adversarial Supervision Makes Layout-to-Image Diffusion Models Thrive
di: Li, Yumeng, et al.
Pubblicazione: (2024)
di: Li, Yumeng, et al.
Pubblicazione: (2024)
Domain-Aware Fine-Tuning of Foundation Models
di: Kaplan, Ugur Ali, et al.
Pubblicazione: (2024)
di: Kaplan, Ugur Ali, et al.
Pubblicazione: (2024)
Multimodal Long Video Modeling Based on Temporal Dynamic Context
di: Hao, Haoran, et al.
Pubblicazione: (2025)
di: Hao, Haoran, et al.
Pubblicazione: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
Video Face Re-Aging: Toward Temporally Consistent Face Re-Aging
di: Muqeet, Abdul, et al.
Pubblicazione: (2023)
di: Muqeet, Abdul, et al.
Pubblicazione: (2023)
STIV: Scalable Text and Image Conditioned Video Generation
di: Lin, Zongyu, et al.
Pubblicazione: (2024)
di: Lin, Zongyu, et al.
Pubblicazione: (2024)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2025)
di: Li, Quanhao, et al.
Pubblicazione: (2025)
LayerT2V: A Unified Multi-Layer Video Generation Framework
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
di: Abbasi, Mehryar, et al.
Pubblicazione: (2024)
di: Abbasi, Mehryar, et al.
Pubblicazione: (2024)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
di: Cai, Minghong, et al.
Pubblicazione: (2024)
di: Cai, Minghong, et al.
Pubblicazione: (2024)
Latent Space Probing for Adult Content Detection in Video Generative Models
di: Khatri, Alizishaan, et al.
Pubblicazione: (2026)
di: Khatri, Alizishaan, et al.
Pubblicazione: (2026)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2026)
di: Li, Quanhao, et al.
Pubblicazione: (2026)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
di: Wang, Zhouxia, et al.
Pubblicazione: (2023)
di: Wang, Zhouxia, et al.
Pubblicazione: (2023)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
di: Yu, Lijun
Pubblicazione: (2024)
di: Yu, Lijun
Pubblicazione: (2024)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
di: Zheng, Sixiao, et al.
Pubblicazione: (2025)
di: Zheng, Sixiao, et al.
Pubblicazione: (2025)
Video-based Music Generation
di: Sulun, Serkan
Pubblicazione: (2026)
di: Sulun, Serkan
Pubblicazione: (2026)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
di: Tang, Shixiang, et al.
Pubblicazione: (2025)
di: Tang, Shixiang, et al.
Pubblicazione: (2025)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
di: Girdhar, Rohit, et al.
Pubblicazione: (2023)
di: Girdhar, Rohit, et al.
Pubblicazione: (2023)
Low-Resolution Object Recognition with Cross-Resolution Relational Contrastive Distillation
di: Zhang, Kangkai, et al.
Pubblicazione: (2024)
di: Zhang, Kangkai, et al.
Pubblicazione: (2024)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
di: Chen, Weifeng, et al.
Pubblicazione: (2023)
di: Chen, Weifeng, et al.
Pubblicazione: (2023)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
di: Liao, Yi, et al.
Pubblicazione: (2024)
di: Liao, Yi, et al.
Pubblicazione: (2024)
Diffusion Model-Based Video Editing: A Survey
di: Sun, Wenhao, et al.
Pubblicazione: (2024)
di: Sun, Wenhao, et al.
Pubblicazione: (2024)
Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
di: Chiu, Pin-Yen, et al.
Pubblicazione: (2025)
di: Chiu, Pin-Yen, et al.
Pubblicazione: (2025)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
di: Yang, Dejie, et al.
Pubblicazione: (2024)
di: Yang, Dejie, et al.
Pubblicazione: (2024)
Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2022)
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2022)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
di: Lin, Ronghao, et al.
Pubblicazione: (2024)
di: Lin, Ronghao, et al.
Pubblicazione: (2024)
MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2025)
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2025)
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
di: Lin, Yan-Bo, et al.
Pubblicazione: (2026)
di: Lin, Yan-Bo, et al.
Pubblicazione: (2026)
Using AI to Summarize US Presidential Campaign TV Advertisement Videos, 1952-2012
di: Breuer, Adam, et al.
Pubblicazione: (2025)
di: Breuer, Adam, et al.
Pubblicazione: (2025)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
di: Chen, Baiyu, et al.
Pubblicazione: (2025)
di: Chen, Baiyu, et al.
Pubblicazione: (2025)
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
di: Wang, Qiang, et al.
Pubblicazione: (2025)
di: Wang, Qiang, et al.
Pubblicazione: (2025)
Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising
di: Singer, Assaf, et al.
Pubblicazione: (2025)
di: Singer, Assaf, et al.
Pubblicazione: (2025)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
di: Zhang, Shiyi, et al.
Pubblicazione: (2026)
di: Zhang, Shiyi, et al.
Pubblicazione: (2026)
Flow Generator Matching
di: Huang, Zemin, et al.
Pubblicazione: (2024)
di: Huang, Zemin, et al.
Pubblicazione: (2024)
Generating Illustrated Instructions
di: Menon, Sachit, et al.
Pubblicazione: (2023)
di: Menon, Sachit, et al.
Pubblicazione: (2023)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
Discover Your Neighbors: Advanced Stable Test-Time Adaptation in Dynamic World
di: Jiang, Qinting, et al.
Pubblicazione: (2024)
di: Jiang, Qinting, et al.
Pubblicazione: (2024)
Improving Visual Representation Alignment Generation with GRPO
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Divide & Bind Your Attention for Improved Generative Semantic Nursing
di: Li, Yumeng, et al.
Pubblicazione: (2023) -
Adversarial Supervision Makes Layout-to-Image Diffusion Models Thrive
di: Li, Yumeng, et al.
Pubblicazione: (2024) -
Domain-Aware Fine-Tuning of Foundation Models
di: Kaplan, Ugur Ali, et al.
Pubblicazione: (2024) -
Multimodal Long Video Modeling Based on Temporal Dynamic Context
di: Hao, Haoran, et al.
Pubblicazione: (2025) -
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)