Image-to-Video Diffusion: From Foundations to Open Frontiers
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xianlong, Pan, Wenbo, Zhou, Shijia, Li, Ke, Wang, Yuqi, Ye, Zeyu, Zhang, Hangtao, Zhang, Leo Yu, Jia, Xiaohua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual-branch Robust Unlearnable Examples
by: Wang, Xianlong, et al.
Published: (2026)
by: Wang, Xianlong, et al.
Published: (2026)
Detector Collapse: Physical-World Backdooring Object Detection to Catastrophic Overload or Blindness in Autonomous Driving
by: Zhang, Hangtao, et al.
Published: (2024)
by: Zhang, Hangtao, et al.
Published: (2024)
PB-UAP: Hybrid Universal Adversarial Attack For Image Segmentation
by: Song, Yufei, et al.
Published: (2024)
by: Song, Yufei, et al.
Published: (2024)
ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
by: Wang, Yichen, et al.
Published: (2025)
by: Wang, Yichen, et al.
Published: (2025)
Unlearnable 3D Point Clouds: Class-wise Transformation Is All You Need
by: Wang, Xianlong, et al.
Published: (2024)
by: Wang, Xianlong, et al.
Published: (2024)
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
by: Zhang, Boqiang, et al.
Published: (2025)
by: Zhang, Boqiang, et al.
Published: (2025)
Diffusion$^2$: Dynamic 3D Content Generation via Score Composition of Video and Multi-view Diffusion Models
by: Yang, Zeyu, et al.
Published: (2024)
by: Yang, Zeyu, et al.
Published: (2024)
Image-to-Image Matching via Foundation Models: A New Perspective for Open-Vocabulary Semantic Segmentation
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
Test-Time Backdoor Detection for Object Detection Models
by: Zhang, Hangtao, et al.
Published: (2025)
by: Zhang, Hangtao, et al.
Published: (2025)
Protective Perturbations against Unauthorized Data Usage in Diffusion-based Image Generation
by: Peng, Sen, et al.
Published: (2024)
by: Peng, Sen, et al.
Published: (2024)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
by: Peng, Bohao, et al.
Published: (2024)
by: Peng, Bohao, et al.
Published: (2024)
Detecting and Corrupting Convolution-based Unlearnable Examples
by: Li, Minghui, et al.
Published: (2023)
by: Li, Minghui, et al.
Published: (2023)
ECLIPSE: Expunging Clean-label Indiscriminate Poisons via Sparse Diffusion Purification
by: Wang, Xianlong, et al.
Published: (2024)
by: Wang, Xianlong, et al.
Published: (2024)
EgoLCD: Egocentric Video Generation with Long Context Diffusion
by: Zhang, Liuzhou, et al.
Published: (2025)
by: Zhang, Liuzhou, et al.
Published: (2025)
VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement
by: Zhang, Shulian, et al.
Published: (2025)
by: Zhang, Shulian, et al.
Published: (2025)
SegTrans: Transferable Adversarial Examples for Segmentation Models
by: Song, Yufei, et al.
Published: (2025)
by: Song, Yufei, et al.
Published: (2025)
DarkHash: A Data-Free Backdoor Attack Against Deep Hashing
by: Zhou, Ziqi, et al.
Published: (2025)
by: Zhou, Ziqi, et al.
Published: (2025)
FaceTracer: Unveiling Source Identities from Swapped Face Images and Videos for Fraud Prevention
by: Zhang, Zhongyi, et al.
Published: (2024)
by: Zhang, Zhongyi, et al.
Published: (2024)
SGCap: Decoding Semantic Group for Zero-shot Video Captioning
by: Pan, Zeyu, et al.
Published: (2025)
by: Pan, Zeyu, et al.
Published: (2025)
Beyond Sliders: Mastering the Art of Diffusion-based Image Manipulation
by: Tang, Yufei, et al.
Published: (2025)
by: Tang, Yufei, et al.
Published: (2025)
Generative Models at the Frontier of Compression: A Survey on Generative Face Video Coding
by: Chen, Bolin, et al.
Published: (2025)
by: Chen, Bolin, et al.
Published: (2025)
LaSe-E2V: Towards Language-guided Semantic-Aware Event-to-Video Reconstruction
by: Chen, Kanghao, et al.
Published: (2024)
by: Chen, Kanghao, et al.
Published: (2024)
Breaking Barriers in Physical-World Adversarial Examples: Improving Robustness and Transferability via Robust Feature
by: Wang, Yichen, et al.
Published: (2024)
by: Wang, Yichen, et al.
Published: (2024)
Image Segmentation in Foundation Model Era: A Survey
by: Zhou, Tianfei, et al.
Published: (2024)
by: Zhou, Tianfei, et al.
Published: (2024)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
by: Su, Zhaochen, et al.
Published: (2025)
by: Su, Zhaochen, et al.
Published: (2025)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
by: Zheng, Duo, et al.
Published: (2024)
by: Zheng, Duo, et al.
Published: (2024)
From Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification
by: Zhang, Ke, et al.
Published: (2026)
by: Zhang, Ke, et al.
Published: (2026)
Enabling Versatile Controls for Video Diffusion Models
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
by: Xu, Tian-Xing, et al.
Published: (2025)
by: Xu, Tian-Xing, et al.
Published: (2025)
Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis
by: Lu, Yanzuo, et al.
Published: (2025)
by: Lu, Yanzuo, et al.
Published: (2025)
Versatile Transition Generation with Image-to-Video Diffusion
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors
by: Bin, Yanrui, et al.
Published: (2025)
by: Bin, Yanrui, et al.
Published: (2025)
MedDiff-FM: A Diffusion-based Foundation Model for Versatile Medical Image Applications
by: Yu, Yongrui, et al.
Published: (2024)
by: Yu, Yongrui, et al.
Published: (2024)
CoD: A Diffusion Foundation Model for Image Compression
by: Jia, Zhaoyang, et al.
Published: (2025)
by: Jia, Zhaoyang, et al.
Published: (2025)
AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset
by: Zhang, Haiyu, et al.
Published: (2025)
by: Zhang, Haiyu, et al.
Published: (2025)
Survey of Video Diffusion Models: Foundations, Implementations, and Applications
by: Wang, Yimu, et al.
Published: (2025)
by: Wang, Yimu, et al.
Published: (2025)
Neuromorphic Synergy for Video Binarization
by: Lin, Shijie, et al.
Published: (2024)
by: Lin, Shijie, et al.
Published: (2024)
OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation Triad
by: Tang, Luyao, et al.
Published: (2025)
by: Tang, Luyao, et al.
Published: (2025)
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
by: Wang, Qilin, et al.
Published: (2024)
by: Wang, Qilin, et al.
Published: (2024)
Similar Items
-
Dual-branch Robust Unlearnable Examples
by: Wang, Xianlong, et al.
Published: (2026) -
Detector Collapse: Physical-World Backdooring Object Detection to Catastrophic Overload or Blindness in Autonomous Driving
by: Zhang, Hangtao, et al.
Published: (2024) -
PB-UAP: Hybrid Universal Adversarial Attack For Image Segmentation
by: Song, Yufei, et al.
Published: (2024) -
ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
by: Wang, Yichen, et al.
Published: (2025) -
Unlearnable 3D Point Clouds: Class-wise Transformation Is All You Need
by: Wang, Xianlong, et al.
Published: (2024)