Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Muyang, Guo, Hanzhong, Lin, Junxiong, Yu, Yizhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neighboring Autoregressive Modeling for Efficient Visual Generation
von: He, Yefei, et al.
Veröffentlicht: (2025)
von: He, Yefei, et al.
Veröffentlicht: (2025)
LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
Bora: Biomedical Generalist Video Generation Model
von: Sun, Weixiang, et al.
Veröffentlicht: (2024)
von: Sun, Weixiang, et al.
Veröffentlicht: (2024)
VEnhancer: Generative Space-Time Enhancement for Video Generation
von: He, Jingwen, et al.
Veröffentlicht: (2024)
von: He, Jingwen, et al.
Veröffentlicht: (2024)
Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms
von: Stojanovski, David, et al.
Veröffentlicht: (2024)
von: Stojanovski, David, et al.
Veröffentlicht: (2024)
Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
MoVideo: Motion-Aware Video Generation with Diffusion Models
von: Liang, Jingyun, et al.
Veröffentlicht: (2023)
von: Liang, Jingyun, et al.
Veröffentlicht: (2023)
Unveiling Hidden Details: A RAW Data-Enhanced Paradigm for Real-World Super-Resolution
von: Peng, Long, et al.
Veröffentlicht: (2024)
von: Peng, Long, et al.
Veröffentlicht: (2024)
Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution
von: Lin, Junxiong, et al.
Veröffentlicht: (2024)
von: Lin, Junxiong, et al.
Veröffentlicht: (2024)
Tractography-Guided Dual-Label Collaborative Learning for Multi-Modal Cranial Nerves Parcellation
von: Xie, Lei, et al.
Veröffentlicht: (2025)
von: Xie, Lei, et al.
Veröffentlicht: (2025)
Improved Video VAE for Latent Video Diffusion Model
von: Wu, Pingyu, et al.
Veröffentlicht: (2024)
von: Wu, Pingyu, et al.
Veröffentlicht: (2024)
SF-V: Single Forward Video Generation Model
von: Zhang, Zhixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhixing, et al.
Veröffentlicht: (2024)
RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space
von: Liang, Jingyun, et al.
Veröffentlicht: (2025)
von: Liang, Jingyun, et al.
Veröffentlicht: (2025)
Generative Latent Video Compression
von: Guo, Zongyu, et al.
Veröffentlicht: (2025)
von: Guo, Zongyu, et al.
Veröffentlicht: (2025)
Taming Diffusion Transformer for Efficient Mobile Video Generation in Seconds
von: Wu, Yushu, et al.
Veröffentlicht: (2025)
von: Wu, Yushu, et al.
Veröffentlicht: (2025)
M3-CVC: Controllable Video Compression with Multimodal Generative Models
von: Wan, Rui, et al.
Veröffentlicht: (2024)
von: Wan, Rui, et al.
Veröffentlicht: (2024)
CMC-Bench: Towards a New Paradigm of Visual Signal Compression
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
von: Chen, Liuhan, et al.
Veröffentlicht: (2024)
von: Chen, Liuhan, et al.
Veröffentlicht: (2024)
Taming Lookup Tables for Efficient Image Retouching
von: Yang, Sidi, et al.
Veröffentlicht: (2024)
von: Yang, Sidi, et al.
Veröffentlicht: (2024)
A Spatio-temporal Aligned SUNet Model for Low-light Video Enhancement
von: Lin, Ruirui, et al.
Veröffentlicht: (2024)
von: Lin, Ruirui, et al.
Veröffentlicht: (2024)
Unsupervised Real-World Super-Resolution via Rectified Flow Degradation Modelling
von: Zhou, Hongyang, et al.
Veröffentlicht: (2025)
von: Zhou, Hongyang, et al.
Veröffentlicht: (2025)
Fed-NDIF: A Noise-Embedded Federated Diffusion Model For Low-Count Whole-Body PET Denoising
von: Zhou, Yinchi, et al.
Veröffentlicht: (2025)
von: Zhou, Yinchi, et al.
Veröffentlicht: (2025)
Overfitting in Histopathology Model Training: The Need for Customized Architectures
von: Alfasly, Saghir, et al.
Veröffentlicht: (2025)
von: Alfasly, Saghir, et al.
Veröffentlicht: (2025)
Learning Phase Distortion with Selective State Space Models for Video Turbulence Mitigation
von: Zhang, Xingguang, et al.
Veröffentlicht: (2025)
von: Zhang, Xingguang, et al.
Veröffentlicht: (2025)
Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation
von: Ge, Xingtong, et al.
Veröffentlicht: (2026)
von: Ge, Xingtong, et al.
Veröffentlicht: (2026)
Efficient Dynamic-NeRF Based Volumetric Video Coding with Rate Distortion Optimization
von: Zhang, Zhiyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiyu, et al.
Veröffentlicht: (2024)
RealisVSR: Detail-enhanced Diffusion for Real-World 4K Video Super-Resolution
von: Zhao, Weisong, et al.
Veröffentlicht: (2025)
von: Zhao, Weisong, et al.
Veröffentlicht: (2025)
How Accurate are Video Quality Models for Diffusion-Based Video Super-Resolution?
von: Herb, Benjamin, et al.
Veröffentlicht: (2026)
von: Herb, Benjamin, et al.
Veröffentlicht: (2026)
Towards Robust Time-of-Flight Depth Denoising with Confidence-Aware Diffusion Model
von: He, Changyong, et al.
Veröffentlicht: (2025)
von: He, Changyong, et al.
Veröffentlicht: (2025)
Towards a General-Purpose Zero-Shot Synthetic Low-Light Image and Video Pipeline
von: Lin, Joanne, et al.
Veröffentlicht: (2025)
von: Lin, Joanne, et al.
Veröffentlicht: (2025)
RepNet-VSR: Reparameterizable Architecture for High-Fidelity Video Super-Resolution
von: Wu, Biao, et al.
Veröffentlicht: (2025)
von: Wu, Biao, et al.
Veröffentlicht: (2025)
Efficient and Accurate Hyperspectral Image Demosaicing with Neural Network Architectures
von: Wisotzky, Eric L., et al.
Veröffentlicht: (2023)
von: Wisotzky, Eric L., et al.
Veröffentlicht: (2023)
A Region of Interest Focused Triple UNet Architecture for Skin Lesion Segmentation
von: Liu, Guoqing, et al.
Veröffentlicht: (2023)
von: Liu, Guoqing, et al.
Veröffentlicht: (2023)
LDMVFI: Video Frame Interpolation with Latent Diffusion Models
von: Danier, Duolikun, et al.
Veröffentlicht: (2023)
von: Danier, Duolikun, et al.
Veröffentlicht: (2023)
Extreme Video Compression with Pre-trained Diffusion Models
von: Li, Bohan, et al.
Veröffentlicht: (2024)
von: Li, Bohan, et al.
Veröffentlicht: (2024)
Dataset Distillation in Medical Imaging: A Feasibility Study
von: Li, Muyang, et al.
Veröffentlicht: (2024)
von: Li, Muyang, et al.
Veröffentlicht: (2024)
PA-SAM: Prompt Adapter SAM for High-Quality Image Segmentation
von: Xie, Zhaozhi, et al.
Veröffentlicht: (2024)
von: Xie, Zhaozhi, et al.
Veröffentlicht: (2024)
Information Prebuilt Recurrent Reconstruction Network for Video Super-Resolution
von: Wang, Shuyun, et al.
Veröffentlicht: (2021)
von: Wang, Shuyun, et al.
Veröffentlicht: (2021)
From General to Specialized: The Need for Foundational Models in Agriculture
von: Nedungadi, Vishal, et al.
Veröffentlicht: (2025)
von: Nedungadi, Vishal, et al.
Veröffentlicht: (2025)
IN2OUT: Fine-Tuning Video Inpainting Model for Video Outpainting Using Hierarchical Discriminator
von: Youn, Sangwoo, et al.
Veröffentlicht: (2025)
von: Youn, Sangwoo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Neighboring Autoregressive Modeling for Efficient Visual Generation
von: He, Yefei, et al.
Veröffentlicht: (2025) -
LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models
von: Cheng, Yu, et al.
Veröffentlicht: (2025) -
Bora: Biomedical Generalist Video Generation Model
von: Sun, Weixiang, et al.
Veröffentlicht: (2024) -
VEnhancer: Generative Space-Time Enhancement for Video Generation
von: He, Jingwen, et al.
Veröffentlicht: (2024) -
Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms
von: Stojanovski, David, et al.
Veröffentlicht: (2024)