OSV: One Step is Enough for High-Quality Image to Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Xiaofeng, Jiang, Zhengkai, Wang, Fu-Yun, Zhang, Jiangning, Chen, Hao, Chi, Mingmin, Wang, Yabiao, Luo, Wenhan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
by: Mao, Xiaofeng, et al.
Published: (2024)
by: Mao, Xiaofeng, et al.
Published: (2024)
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
by: Wang, Qilin, et al.
Published: (2024)
by: Wang, Qilin, et al.
Published: (2024)
AdapNet: Adaptive Noise-Based Network for Low-Quality Image Retrieval
by: Zhang, Sihe, et al.
Published: (2024)
by: Zhang, Sihe, et al.
Published: (2024)
PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
by: Pan, Yanjie, et al.
Published: (2025)
by: Pan, Yanjie, et al.
Published: (2025)
Once Is Enough: Lightweight DiT-Based Video Virtual Try-On via One-Time Garment Appearance Injection
by: Pan, Yanjie, et al.
Published: (2025)
by: Pan, Yanjie, et al.
Published: (2025)
PVG: Progressive Vision Graph for Vision Recognition
by: Wu, Jiafu, et al.
Published: (2023)
by: Wu, Jiafu, et al.
Published: (2023)
SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment
by: Sun, Yanxiao, et al.
Published: (2025)
by: Sun, Yanxiao, et al.
Published: (2025)
InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation
by: Liu, Xingchao, et al.
Published: (2023)
by: Liu, Xingchao, et al.
Published: (2023)
Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary Detection
by: Wang, Haoxuan, et al.
Published: (2024)
by: Wang, Haoxuan, et al.
Published: (2024)
Learning Unified Reference Representation for Unsupervised Multi-class Anomaly Detection
by: He, Liren, et al.
Published: (2024)
by: He, Liren, et al.
Published: (2024)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
by: Xue, Zhucun, et al.
Published: (2025)
by: Xue, Zhucun, et al.
Published: (2025)
Dual-Interrelated Diffusion Model for Few-Shot Anomaly Image Generation
by: Jin, Ying, et al.
Published: (2024)
by: Jin, Ying, et al.
Published: (2024)
Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory
by: Liu, Jinzhuo, et al.
Published: (2026)
by: Liu, Jinzhuo, et al.
Published: (2026)
MotionMaster: Training-free Camera Motion Transfer For Video Generation
by: Hu, Teng, et al.
Published: (2024)
by: Hu, Teng, et al.
Published: (2024)
UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
by: Wang, Hanyang, et al.
Published: (2025)
by: Wang, Hanyang, et al.
Published: (2025)
AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model
by: Hu, Teng, et al.
Published: (2023)
by: Hu, Teng, et al.
Published: (2023)
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$
by: Zhang, Jiangning, et al.
Published: (2025)
by: Zhang, Jiangning, et al.
Published: (2025)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
by: Mao, Xiaofeng, et al.
Published: (2026)
by: Mao, Xiaofeng, et al.
Published: (2026)
SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation
by: Hu, Teng, et al.
Published: (2024)
by: Hu, Teng, et al.
Published: (2024)
PiT: Progressive Diffusion Transformer
by: Wu, Jiafu, et al.
Published: (2025)
by: Wu, Jiafu, et al.
Published: (2025)
VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement
by: Zhang, Shulian, et al.
Published: (2025)
by: Zhang, Shulian, et al.
Published: (2025)
PointSeg: A Training-Free Paradigm for 3D Scene Segmentation via Foundation Models
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion Generation
by: Wang, Yabiao, et al.
Published: (2024)
by: Wang, Yabiao, et al.
Published: (2024)
Towards One-step Causal Video Generation via Adversarial Self-Distillation
by: Yang, Yongqi, et al.
Published: (2025)
by: Yang, Yongqi, et al.
Published: (2025)
Leveraging Fine-Grained Information and Noise Decoupling for Remote Sensing Change Detection
by: Du, Qiangang, et al.
Published: (2024)
by: Du, Qiangang, et al.
Published: (2024)
MARRS: Masked Autoregressive Unit-based Reaction Synthesis
by: Wang, Yabiao, et al.
Published: (2025)
by: Wang, Yabiao, et al.
Published: (2025)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
by: Surkov, Viacheslav, et al.
Published: (2024)
by: Surkov, Viacheslav, et al.
Published: (2024)
Foundation Cures Personalization: Improving Personalized Models' Prompt Consistency via Hidden Foundation Knowledge
by: Cai, Yiyang, et al.
Published: (2024)
by: Cai, Yiyang, et al.
Published: (2024)
Exploring Real&Synthetic Dataset and Linear Attention in Image Restoration
by: Du, Yuzhen, et al.
Published: (2024)
by: Du, Yuzhen, et al.
Published: (2024)
VI3DRM:Towards meticulous 3D Reconstruction from Sparse Views via Photo-Realistic Novel View Synthesis
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
by: Cai, Yuxuan, et al.
Published: (2025)
by: Cai, Yuxuan, et al.
Published: (2025)
UnSeg: One Universal Unlearnable Example Generator is Enough against All Image Segmentation
by: Sun, Ye, et al.
Published: (2024)
by: Sun, Ye, et al.
Published: (2024)
Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
UltraGen: High-Resolution Video Generation with Hierarchical Attention
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Similar Items
-
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
by: Mao, Xiaofeng, et al.
Published: (2024) -
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
by: Wang, Qilin, et al.
Published: (2024) -
AdapNet: Adaptive Noise-Based Network for Low-Quality Image Retrieval
by: Zhang, Sihe, et al.
Published: (2024) -
PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
by: Pan, Yanjie, et al.
Published: (2025) -
Once Is Enough: Lightweight DiT-Based Video Virtual Try-On via One-Time Garment Appearance Injection
by: Pan, Yanjie, et al.
Published: (2025)