Mojito: Motion Trajectory and Intensity Control for Video Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | He, Xuehai, Wang, Shuohang, Yang, Jianwei, Wu, Xiaoxia, Wang, Yiping, Wang, Kuan, Zhan, Zheng, Ruwase, Olatunji, Shen, Yelong, Wang, Xin Eric |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
por: Wang, Yiping, et al.
Publicado: (2024)
por: Wang, Yiping, et al.
Publicado: (2024)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
por: Zhang, Zhen, et al.
Publicado: (2025)
por: Zhang, Zhen, et al.
Publicado: (2025)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
por: Huang, Zeyi, et al.
Publicado: (2026)
por: Huang, Zeyi, et al.
Publicado: (2026)
FastPersist: Accelerating Model Checkpointing in Deep Learning
por: Wang, Guanhua, et al.
Publicado: (2024)
por: Wang, Guanhua, et al.
Publicado: (2024)
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
por: Wang, Guanhua, et al.
Publicado: (2024)
por: Wang, Guanhua, et al.
Publicado: (2024)
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
por: Zheng, Kaizhi, et al.
Publicado: (2023)
por: Zheng, Kaizhi, et al.
Publicado: (2023)
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
por: Wang, Yiping, et al.
Publicado: (2025)
por: Wang, Yiping, et al.
Publicado: (2025)
Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional Encoding
por: Zhang, Zhenyu, et al.
Publicado: (2024)
por: Zhang, Zhenyu, et al.
Publicado: (2024)
Temperature-Centric Investigation of Speculative Decoding with Knowledge Distillation
por: Ouyang, Siru, et al.
Publicado: (2024)
por: Ouyang, Siru, et al.
Publicado: (2024)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
por: Zhang, Rongzhi, et al.
Publicado: (2024)
por: Zhang, Rongzhi, et al.
Publicado: (2024)
Motion Prompting: Controlling Video Generation with Motion Trajectories
por: Geng, Daniel, et al.
Publicado: (2024)
por: Geng, Daniel, et al.
Publicado: (2024)
Mojito: LLM-Aided Motion Instructor with Jitter-Reduced Inertial Tokens
por: Shan, Ziwei, et al.
Publicado: (2025)
por: Shan, Ziwei, et al.
Publicado: (2025)
MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
por: Lei, Guojun, et al.
Publicado: (2025)
por: Lei, Guojun, et al.
Publicado: (2025)
Multi-LoRA Composition for Image Generation
por: Zhong, Ming, et al.
Publicado: (2024)
por: Zhong, Ming, et al.
Publicado: (2024)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
por: Lian, Xinyu, et al.
Publicado: (2025)
por: Lian, Xinyu, et al.
Publicado: (2025)
ThetaEvolve: Test-time Learning on Open Problems
por: Wang, Yiping, et al.
Publicado: (2025)
por: Wang, Yiping, et al.
Publicado: (2025)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
por: Gupta, Ahan, et al.
Publicado: (2026)
por: Gupta, Ahan, et al.
Publicado: (2026)
ComCLIP: Training-Free Compositional Image and Text Matching
por: Jiang, Kenan, et al.
Publicado: (2022)
por: Jiang, Kenan, et al.
Publicado: (2022)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
por: Li, Quanhao, et al.
Publicado: (2025)
por: Li, Quanhao, et al.
Publicado: (2025)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
por: Tanaka, Masahiro, et al.
Publicado: (2025)
por: Tanaka, Masahiro, et al.
Publicado: (2025)
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
por: Shi, Shuwei, et al.
Publicado: (2024)
por: Shi, Shuwei, et al.
Publicado: (2024)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
por: Li, Quanhao, et al.
Publicado: (2026)
por: Li, Quanhao, et al.
Publicado: (2026)
MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
por: He, Xuehai, et al.
Publicado: (2025)
por: He, Xuehai, et al.
Publicado: (2025)
Self-Evolving 3D Scene Generation from a Single Image
por: Zheng, Kaizhi, et al.
Publicado: (2025)
por: Zheng, Kaizhi, et al.
Publicado: (2025)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
por: Zhu, Chenhui, et al.
Publicado: (2025)
por: Zhu, Chenhui, et al.
Publicado: (2025)
OmniParser for Pure Vision Based GUI Agent
por: Lu, Yadong, et al.
Publicado: (2024)
por: Lu, Yadong, et al.
Publicado: (2024)
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
por: Ren, Liliang, et al.
Publicado: (2025)
por: Ren, Liliang, et al.
Publicado: (2025)
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
por: Wang, Boyuan, et al.
Publicado: (2025)
por: Wang, Boyuan, et al.
Publicado: (2025)
Adapting LLM Agents with Universal Feedback in Communication
por: Wang, Kuan, et al.
Publicado: (2023)
por: Wang, Kuan, et al.
Publicado: (2023)
FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation
por: He, Xuehai, et al.
Publicado: (2024)
por: He, Xuehai, et al.
Publicado: (2024)
TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
por: Wang, Yufu, et al.
Publicado: (2024)
por: Wang, Yufu, et al.
Publicado: (2024)
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
por: Zheng, Kaizhi, et al.
Publicado: (2022)
por: Zheng, Kaizhi, et al.
Publicado: (2022)
MotionClone: Training-Free Motion Cloning for Controllable Video Generation
por: Ling, Pengyang, et al.
Publicado: (2024)
por: Ling, Pengyang, et al.
Publicado: (2024)
FineXtrol: Controllable Motion Generation via Fine-Grained Text
por: Shen, Keming, et al.
Publicado: (2025)
por: Shen, Keming, et al.
Publicado: (2025)
Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning
por: Chen, Yifang, et al.
Publicado: (2024)
por: Chen, Yifang, et al.
Publicado: (2024)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
por: He, Xuehai, et al.
Publicado: (2024)
por: He, Xuehai, et al.
Publicado: (2024)
From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation
por: Yang, Fan, et al.
Publicado: (2025)
por: Yang, Fan, et al.
Publicado: (2025)
Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
por: Chu, Ruihang, et al.
Publicado: (2025)
por: Chu, Ruihang, et al.
Publicado: (2025)
ReCamDriving: LiDAR-Free Camera-Controlled Novel Trajectory Video Generation
por: Li, Yaokun, et al.
Publicado: (2025)
por: Li, Yaokun, et al.
Publicado: (2025)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
por: Wang, Zhouxia, et al.
Publicado: (2023)
por: Wang, Zhouxia, et al.
Publicado: (2023)
Ejemplares similares
-
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
por: Wang, Yiping, et al.
Publicado: (2024) -
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
por: Zhang, Zhen, et al.
Publicado: (2025) -
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
por: Huang, Zeyi, et al.
Publicado: (2026) -
FastPersist: Accelerating Model Checkpointing in Deep Learning
por: Wang, Guanhua, et al.
Publicado: (2024) -
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
por: Wang, Guanhua, et al.
Publicado: (2024)