Seer: Language Instructed Video Prediction with Latent Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Xianfan, Wen, Chuan, Ye, Weirui, Song, Jiaming, Gao, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
by: Ding, Zheng, et al.
Published: (2025)
by: Ding, Zheng, et al.
Published: (2025)
From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
Nodule-Aligned Latent Space Learning with LLM-Driven Multimodal Diffusion for Lung Nodule Progression Prediction
by: Song, James, et al.
Published: (2026)
by: Song, James, et al.
Published: (2026)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
by: Zhang, David Junhao, et al.
Published: (2023)
by: Zhang, David Junhao, et al.
Published: (2023)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
Loom: Diffusion-Transformer for Interleaved Generation
by: Ye, Mingcheng, et al.
Published: (2025)
by: Ye, Mingcheng, et al.
Published: (2025)
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
by: Liu, Shizhan, et al.
Published: (2025)
by: Liu, Shizhan, et al.
Published: (2025)
LVMark: Robust Watermark for Latent Video Diffusion Models
by: Jang, MinHyuk, et al.
Published: (2024)
by: Jang, MinHyuk, et al.
Published: (2024)
Motion-aware Latent Diffusion Models for Video Frame Interpolation
by: Huang, Zhilin, et al.
Published: (2024)
by: Huang, Zhilin, et al.
Published: (2024)
Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion Model
by: Zhou, Hang, et al.
Published: (2024)
by: Zhou, Hang, et al.
Published: (2024)
Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
by: Lu, Tianyi, et al.
Published: (2023)
by: Lu, Tianyi, et al.
Published: (2023)
Progressive Text-to-Image Diffusion with Soft Latent Direction
by: Ye, YuTeng, et al.
Published: (2023)
by: Ye, YuTeng, et al.
Published: (2023)
Multi-Condition Latent Diffusion Network for Scene-Aware Neural Human Motion Prediction
by: Gao, Xuehao, et al.
Published: (2024)
by: Gao, Xuehao, et al.
Published: (2024)
Plan-X: Instruct Video Generation via Semantic Planning
by: Huang, Lun, et al.
Published: (2025)
by: Huang, Lun, et al.
Published: (2025)
LTX-Video: Realtime Video Latent Diffusion
by: HaCohen, Yoav, et al.
Published: (2024)
by: HaCohen, Yoav, et al.
Published: (2024)
Video Generation with Predictive Latents
by: Zhao, Yian, et al.
Published: (2026)
by: Zhao, Yian, et al.
Published: (2026)
Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
by: Gu, Jiatao, et al.
Published: (2024)
by: Gu, Jiatao, et al.
Published: (2024)
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
by: Zheng, Haojie, et al.
Published: (2026)
by: Zheng, Haojie, et al.
Published: (2026)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)
by: Li, Zongjian, et al.
Published: (2024)
LatentColorization: Latent Diffusion-Based Speaker Video Colorization
by: Ward, Rory, et al.
Published: (2024)
by: Ward, Rory, et al.
Published: (2024)
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
Improved Video VAE for Latent Video Diffusion Model
by: Wu, Pingyu, et al.
Published: (2024)
by: Wu, Pingyu, et al.
Published: (2024)
InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
by: Lin, Kun-Hsiang, et al.
Published: (2025)
by: Lin, Kun-Hsiang, et al.
Published: (2025)
STDiff: Spatio-temporal Diffusion for Continuous Stochastic Video Prediction
by: Ye, Xi, et al.
Published: (2023)
by: Ye, Xi, et al.
Published: (2023)
InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
by: Gan, Yulu, et al.
Published: (2023)
by: Gan, Yulu, et al.
Published: (2023)
Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic Optimization
by: Chang, Yuanyuan, et al.
Published: (2025)
by: Chang, Yuanyuan, et al.
Published: (2025)
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
by: Zhao, Fufangchen, et al.
Published: (2025)
by: Zhao, Fufangchen, et al.
Published: (2025)
Unified Dense Prediction of Video Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
Latte: Latent Diffusion Transformer for Video Generation
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
InstructEngine: Instruction-driven Text-to-Image Alignment
by: Lu, Xingyu, et al.
Published: (2025)
by: Lu, Xingyu, et al.
Published: (2025)
FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space
by: FSVideo Team, et al.
Published: (2026)
by: FSVideo Team, et al.
Published: (2026)
AICL: Action In-Context Learning for Video Diffusion Model
by: Liu, Jianzhi, et al.
Published: (2024)
by: Liu, Jianzhi, et al.
Published: (2024)
Localized Control in Diffusion Models via Latent Vector Prediction
by: Domingo-Gregorio, Pablo, et al.
Published: (2026)
by: Domingo-Gregorio, Pablo, et al.
Published: (2026)
InstructVEdit: A Holistic Approach for Instructional Video Editing
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation
by: Hong, Jihwan, et al.
Published: (2026)
by: Hong, Jihwan, et al.
Published: (2026)
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
SmokeSeer: 3D Gaussian Splatting for Smoke Removal and Scene Reconstruction
by: Jain, Neham, et al.
Published: (2025)
by: Jain, Neham, et al.
Published: (2025)
Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model
by: Mao, Qi, et al.
Published: (2025)
by: Mao, Qi, et al.
Published: (2025)
Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model
by: Peng, Jihua, et al.
Published: (2025)
by: Peng, Jihua, et al.
Published: (2025)
Stable-Makeup: When Real-World Makeup Transfer Meets Diffusion Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
Similar Items
-
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
by: Ding, Zheng, et al.
Published: (2025) -
From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
by: Liu, Jiacheng, et al.
Published: (2025) -
Nodule-Aligned Latent Space Learning with LLM-Driven Multimodal Diffusion for Lung Nodule Progression Prediction
by: Song, James, et al.
Published: (2026) -
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
by: Zhang, David Junhao, et al.
Published: (2023) -
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
by: Wei, Cong, et al.
Published: (2024)