REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion Latents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tian, Rui, Dai, Qi, Bao, Jianmin, Qiu, Kai, Yang, Yifan, Luo, Chong, Wu, Zuxuan, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
von: Tu, Shuyuan, et al.
Veröffentlicht: (2024)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2024)
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
von: Xu, Ziqiang, et al.
Veröffentlicht: (2025)
von: Xu, Ziqiang, et al.
Veröffentlicht: (2025)
FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
von: Li, Quanhao, et al.
Veröffentlicht: (2025)
von: Li, Quanhao, et al.
Veröffentlicht: (2025)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
von: Xing, Zhen, et al.
Veröffentlicht: (2024)
von: Xing, Zhen, et al.
Veröffentlicht: (2024)
MageBench: Bridging Large Multimodal Models to Agents
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
von: Li, Quanhao, et al.
Veröffentlicht: (2026)
von: Li, Quanhao, et al.
Veröffentlicht: (2026)
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
von: Hu, Yaosi, et al.
Veröffentlicht: (2023)
von: Hu, Yaosi, et al.
Veröffentlicht: (2023)
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
OmniVid: A Generative Framework for Universal Video Understanding
von: Wang, Junke, et al.
Veröffentlicht: (2024)
von: Wang, Junke, et al.
Veröffentlicht: (2024)
StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
GlobalPaint: Spatiotemporal Coherent Video Outpainting with Global Feature Guidance
von: Pan, Yueming, et al.
Veröffentlicht: (2026)
von: Pan, Yueming, et al.
Veröffentlicht: (2026)
DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data
von: Jin, Wonjoon, et al.
Veröffentlicht: (2026)
von: Jin, Wonjoon, et al.
Veröffentlicht: (2026)
OmniTracker: Unifying Object Tracking by Tracking-with-Detection
von: Wang, Junke, et al.
Veröffentlicht: (2023)
von: Wang, Junke, et al.
Veröffentlicht: (2023)
A Survey on Video Diffusion Models
von: Xing, Zhen, et al.
Veröffentlicht: (2023)
von: Xing, Zhen, et al.
Veröffentlicht: (2023)
HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
von: Zhou, Ziqin, et al.
Veröffentlicht: (2025)
von: Zhou, Ziqin, et al.
Veröffentlicht: (2025)
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
von: Zhou, Ziwei, et al.
Veröffentlicht: (2026)
von: Zhou, Ziwei, et al.
Veröffentlicht: (2026)
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
von: Chen, Yitong, et al.
Veröffentlicht: (2026)
von: Chen, Yitong, et al.
Veröffentlicht: (2026)
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
von: Meng, Lingchen, et al.
Veröffentlicht: (2024)
von: Meng, Lingchen, et al.
Veröffentlicht: (2024)
LVC-LGMC: Joint Local and Global Motion Compensation for Learned Video Compression
von: Jiang, Wei, et al.
Veröffentlicht: (2024)
von: Jiang, Wei, et al.
Veröffentlicht: (2024)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
StableAnimator: High-Quality Identity-Preserving Human Image Animation
von: Tu, Shuyuan, et al.
Veröffentlicht: (2024)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2024)
Towards Extreme Image Compression with Latent Feature Guidance and Diffusion Prior
von: Li, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Li, Zhiyuan, et al.
Veröffentlicht: (2024)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
GenRec: Unifying Video Generation and Recognition with Diffusion Models
von: Weng, Zejia, et al.
Veröffentlicht: (2024)
von: Weng, Zejia, et al.
Veröffentlicht: (2024)
ArcFlow: Unleashing 2-Step Text-to-Image Generation via High-Precision Non-Linear Flow Distillation
von: Yang, Zihan, et al.
Veröffentlicht: (2026)
von: Yang, Zihan, et al.
Veröffentlicht: (2026)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation
von: Wang, Yanhui, et al.
Veröffentlicht: (2023)
von: Wang, Yanhui, et al.
Veröffentlicht: (2023)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
von: Wang, Junke, et al.
Veröffentlicht: (2024)
von: Wang, Junke, et al.
Veröffentlicht: (2024)
Learning Accurate Segmentation Purely from Self-Supervision
von: You, Zuyao, et al.
Veröffentlicht: (2026)
von: You, Zuyao, et al.
Veröffentlicht: (2026)
Generative Latent Video Compression
von: Guo, Zongyu, et al.
Veröffentlicht: (2025)
von: Guo, Zongyu, et al.
Veröffentlicht: (2025)
Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
von: Lu, Tianyi, et al.
Veröffentlicht: (2023)
von: Lu, Tianyi, et al.
Veröffentlicht: (2023)
Fast Autoregressive Models for Continuous Latent Generation
von: Hang, Tiankai, et al.
Veröffentlicht: (2025)
von: Hang, Tiankai, et al.
Veröffentlicht: (2025)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
CCEdit: Creative and Controllable Video Editing via Diffusion Models
von: Feng, Ruoyu, et al.
Veröffentlicht: (2023)
von: Feng, Ruoyu, et al.
Veröffentlicht: (2023)
DLF: Extreme Image Compression with Dual-generative Latent Fusion
von: Xue, Naifu, et al.
Veröffentlicht: (2025)
von: Xue, Naifu, et al.
Veröffentlicht: (2025)
Latent Spatiotemporal Adaptation for Generalized Face Forgery Video Detection
von: Zhang, Daichi, et al.
Veröffentlicht: (2023)
von: Zhang, Daichi, et al.
Veröffentlicht: (2023)
FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis
von: Jin, Wonjoon, et al.
Veröffentlicht: (2025)
von: Jin, Wonjoon, et al.
Veröffentlicht: (2025)
Generative Latent Coding for Ultra-Low Bitrate Image and Video Compression
von: Qi, Linfeng, et al.
Veröffentlicht: (2025)
von: Qi, Linfeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
von: Tu, Shuyuan, et al.
Veröffentlicht: (2024) -
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
von: Xu, Ziqiang, et al.
Veröffentlicht: (2025) -
FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025) -
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
von: Li, Quanhao, et al.
Veröffentlicht: (2025) -
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
von: Xing, Zhen, et al.
Veröffentlicht: (2024)