Video Generation with Predictive Latents
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yian, Wang, Feng, Guo, Qiushan, Liu, Chang, Ji, Xiangyang, Zhang, Jian, Chen, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
by: Liao, Zijun, et al.
Published: (2025)
by: Liao, Zijun, et al.
Published: (2025)
Comp-Attn: Present-and-Align Attention for Compositional Video Generation
by: Zhang, Hongyu, et al.
Published: (2025)
by: Zhang, Hongyu, et al.
Published: (2025)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability
by: Zhao, Bingchen, et al.
Published: (2026)
by: Zhao, Bingchen, et al.
Published: (2026)
FaceChain-SuDe: Building Derived Class to Inherit Category Attributes for One-shot Subject-Driven Generation
by: Qiao, Pengchong, et al.
Published: (2024)
by: Qiao, Pengchong, et al.
Published: (2024)
PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
by: Du, Shian, et al.
Published: (2025)
by: Du, Shian, et al.
Published: (2025)
iSegMan: Interactive Segment-and-Manipulate 3D Gaussians
by: Zhao, Yian, et al.
Published: (2025)
by: Zhao, Yian, et al.
Published: (2025)
Continuous Latent Diffusion Language Model
by: Guo, Hongcan, et al.
Published: (2026)
by: Guo, Hongcan, et al.
Published: (2026)
UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
by: Du, Shian, et al.
Published: (2025)
by: Du, Shian, et al.
Published: (2025)
Latte: Latent Diffusion Transformer for Video Generation
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
by: Xue, Zeyue, et al.
Published: (2023)
by: Xue, Zeyue, et al.
Published: (2023)
Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
by: Chu, Ruihang, et al.
Published: (2025)
by: Chu, Ruihang, et al.
Published: (2025)
Generative Video Compression with One-Dimensional Latent Representation
by: Zheng, Zihan, et al.
Published: (2026)
by: Zheng, Zihan, et al.
Published: (2026)
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025)
by: Xue, Zeyue, et al.
Published: (2025)
Video Generation Models Are Good Latent Reward Models
by: Mi, Xiaoyue, et al.
Published: (2025)
by: Mi, Xiaoyue, et al.
Published: (2025)
DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning
by: Wei, Yujie, et al.
Published: (2026)
by: Wei, Yujie, et al.
Published: (2026)
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
by: Luo, Xiangyang, et al.
Published: (2026)
by: Luo, Xiangyang, et al.
Published: (2026)
GraCo: Granularity-Controllable Interactive Segmentation
by: Zhao, Yian, et al.
Published: (2024)
by: Zhao, Yian, et al.
Published: (2024)
VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models
by: Zhang, Yabo, et al.
Published: (2024)
by: Zhang, Yabo, et al.
Published: (2024)
LuciBot: Automated Robot Policy Learning from Generated Videos
by: Qiu, Xiaowen, et al.
Published: (2025)
by: Qiu, Xiaowen, et al.
Published: (2025)
Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video
by: Shi, Yahao, et al.
Published: (2025)
by: Shi, Yahao, et al.
Published: (2025)
Automated Label Unification for Multi-Dataset Semantic Segmentation with GNNs
by: Ma, Rong, et al.
Published: (2024)
by: Ma, Rong, et al.
Published: (2024)
Tune-Your-Style: Intensity-tunable 3D Style Transfer with Gaussian Splatting
by: Zhao, Yian, et al.
Published: (2026)
by: Zhao, Yian, et al.
Published: (2026)
Causality-inspired Latent Feature Augmentation for Single Domain Generalization
by: Xu, Jian, et al.
Published: (2024)
by: Xu, Jian, et al.
Published: (2024)
Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation
by: Chen, Hanlin, et al.
Published: (2026)
by: Chen, Hanlin, et al.
Published: (2026)
MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
by: Sun, Mingzhen, et al.
Published: (2024)
by: Sun, Mingzhen, et al.
Published: (2024)
ParCo: Part-Coordinating Text-to-Motion Synthesis
by: Zou, Qiran, et al.
Published: (2024)
by: Zou, Qiran, et al.
Published: (2024)
LEO: Generative Latent Image Animator for Human Video Synthesis
by: Wang, Yaohui, et al.
Published: (2023)
by: Wang, Yaohui, et al.
Published: (2023)
BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
Video Compression Meets Video Generation: Latent Inter-Frame Pruning with Attention Recovery
by: Menn, Dennis, et al.
Published: (2026)
by: Menn, Dennis, et al.
Published: (2026)
Generative Latent Video Compression
by: Guo, Zongyu, et al.
Published: (2025)
by: Guo, Zongyu, et al.
Published: (2025)
RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention
by: Chen, Zhan, et al.
Published: (2025)
by: Chen, Zhan, et al.
Published: (2025)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
by: Chen, Harold Haodong, et al.
Published: (2026)
by: Chen, Harold Haodong, et al.
Published: (2026)
Cross-Domain Vessel Segmentation via Latent Similarity Mining and Iterative Co-Optimization
by: Guo, Zhanqiang, et al.
Published: (2026)
by: Guo, Zhanqiang, et al.
Published: (2026)
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
by: Lv, Wenyu, et al.
Published: (2024)
by: Lv, Wenyu, et al.
Published: (2024)
Latent Video Dataset Distillation
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
by: Zhang, David Junhao, et al.
Published: (2023)
by: Zhang, David Junhao, et al.
Published: (2023)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
by: Liu, Henglin, et al.
Published: (2025)
by: Liu, Henglin, et al.
Published: (2025)
Similar Items
-
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
by: Wang, Feng, et al.
Published: (2026) -
RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
by: Liao, Zijun, et al.
Published: (2025) -
Comp-Attn: Present-and-Align Attention for Compositional Video Generation
by: Zhang, Hongyu, et al.
Published: (2025) -
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
by: Liu, Zhiheng, et al.
Published: (2025) -
Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability
by: Zhao, Bingchen, et al.
Published: (2026)