SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, YaoYang, Zhang, Yuechen, Li, Wenbo, Zhao, Yufei, Liu, Rui, Chen, Long |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained Guidance
by: Shi, Shuwei, et al.
Published: (2024)
by: Shi, Shuwei, et al.
Published: (2024)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
by: Peng, Bohao, et al.
Published: (2024)
by: Peng, Bohao, et al.
Published: (2024)
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
by: Yao, Yuan, et al.
Published: (2025)
by: Yao, Yuan, et al.
Published: (2025)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
by: Zhang, Shuai, et al.
Published: (2025)
by: Zhang, Shuai, et al.
Published: (2025)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025)
by: Qiu, Haonan, et al.
Published: (2025)
V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction
by: Zhao, Yiming, et al.
Published: (2025)
by: Zhao, Yiming, et al.
Published: (2025)
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
Training-Free Efficient Video Generation via Dynamic Token Carving
by: Zhang, Yuechen, et al.
Published: (2025)
by: Zhang, Yuechen, et al.
Published: (2025)
Generative Inbetweening through Frame-wise Conditions-Driven Video Generation
by: Zhu, Tianyi, et al.
Published: (2024)
by: Zhu, Tianyi, et al.
Published: (2024)
Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention
by: Li, Peng, et al.
Published: (2024)
by: Li, Peng, et al.
Published: (2024)
FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation
by: Ge, Yunyang, et al.
Published: (2025)
by: Ge, Yunyang, et al.
Published: (2025)
V-Bridge: Bridging Video Generative Priors to Versatile Few-shot Image Restoration
by: Zheng, Shenghe, et al.
Published: (2026)
by: Zheng, Shenghe, et al.
Published: (2026)
Pushing the Boundaries of State Space Models for Image and Video Generation
by: Hong, Yicong, et al.
Published: (2025)
by: Hong, Yicong, et al.
Published: (2025)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
by: Yang, Haibo, et al.
Published: (2024)
by: Yang, Haibo, et al.
Published: (2024)
PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
by: Du, Shian, et al.
Published: (2025)
by: Du, Shian, et al.
Published: (2025)
MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers
by: Zhang, Yuechen, et al.
Published: (2025)
by: Zhang, Yuechen, et al.
Published: (2025)
Sparse Refinement for Efficient High-Resolution Semantic Segmentation
by: Liu, Zhijian, et al.
Published: (2024)
by: Liu, Zhijian, et al.
Published: (2024)
Swift Parameter-free Attention Network for Efficient Super-Resolution
by: Wan, Cheng, et al.
Published: (2023)
by: Wan, Cheng, et al.
Published: (2023)
ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices
by: Du, Kunpeng, et al.
Published: (2026)
by: Du, Kunpeng, et al.
Published: (2026)
FlashVideo: A Framework for Swift Inference in Text-to-Video Generation
by: Lei, Bin, et al.
Published: (2023)
by: Lei, Bin, et al.
Published: (2023)
HRSAM: Efficient Interactive Segmentation in High-Resolution Images
by: Huang, You, et al.
Published: (2024)
by: Huang, You, et al.
Published: (2024)
SwiftTailor: Efficient 3D Garment Generation with Geometry Image Representation
by: Pham, Phuc, et al.
Published: (2026)
by: Pham, Phuc, et al.
Published: (2026)
Enhancing High-Resolution 3D Generation through Pixel-wise Gradient Clipping
by: Pan, Zijie, et al.
Published: (2023)
by: Pan, Zijie, et al.
Published: (2023)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
by: Ren, Weiming, et al.
Published: (2024)
by: Ren, Weiming, et al.
Published: (2024)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
by: Guo, Xun, et al.
Published: (2023)
by: Guo, Xun, et al.
Published: (2023)
High Quality Segmentation for Ultra High-resolution Images
by: Shen, Tiancheng, et al.
Published: (2021)
by: Shen, Tiancheng, et al.
Published: (2021)
Bilateral Reference for High-Resolution Dichotomous Image Segmentation
by: Zheng, Peng, et al.
Published: (2024)
by: Zheng, Peng, et al.
Published: (2024)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
Promoting Segment Anything Model towards Highly Accurate Dichotomous Image Segmentation
by: Liu, Xianjie, et al.
Published: (2023)
by: Liu, Xianjie, et al.
Published: (2023)
Progressive Image Restoration via Text-Conditioned Video Generation
by: Kang, Peng, et al.
Published: (2025)
by: Kang, Peng, et al.
Published: (2025)
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
by: Cai, Qi, et al.
Published: (2025)
by: Cai, Qi, et al.
Published: (2025)
AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling
by: Mai, Ziyang, et al.
Published: (2026)
by: Mai, Ziyang, et al.
Published: (2026)
SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation
by: Ma, Yingzi, et al.
Published: (2026)
by: Ma, Yingzi, et al.
Published: (2026)
UNetMamba: An Efficient UNet-Like Mamba for Semantic Segmentation of High-Resolution Remote Sensing Images
by: Zhu, Enze, et al.
Published: (2024)
by: Zhu, Enze, et al.
Published: (2024)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
by: Zhang, Zhongwei, et al.
Published: (2025)
by: Zhang, Zhongwei, et al.
Published: (2025)
DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis
by: Teng, Yao, et al.
Published: (2024)
by: Teng, Yao, et al.
Published: (2024)
TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
by: Li, Pengxiang, et al.
Published: (2023)
by: Li, Pengxiang, et al.
Published: (2023)
HRDecoder: High-Resolution Decoder Network for Fundus Image Lesion Segmentation
by: Ding, Ziyuan, et al.
Published: (2024)
by: Ding, Ziyuan, et al.
Published: (2024)
TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Similar Items
-
ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained Guidance
by: Shi, Shuwei, et al.
Published: (2024) -
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
by: Peng, Bohao, et al.
Published: (2024) -
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
by: Yao, Yuan, et al.
Published: (2025) -
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
by: Zhang, Shilong, et al.
Published: (2025) -
MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
by: Zhang, Shuai, et al.
Published: (2025)