FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Ge, Yunyang, Cheng, Xinhua, Zhao, Chengshu, He, Xianyi, Yuan, Shenghai, Lin, Bin, Zhu, Bin, Yuan, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning
by: Ge, Yunyang, et al.
Published: (2026)
by: Ge, Yunyang, et al.
Published: (2026)
SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video
by: Zhao, Chengshu, et al.
Published: (2025)
by: Zhao, Chengshu, et al.
Published: (2025)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
by: Lin, Bin, et al.
Published: (2025)
by: Lin, Bin, et al.
Published: (2025)
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
by: Chen, Liuhan, et al.
Published: (2024)
by: Chen, Liuhan, et al.
Published: (2024)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)
by: Li, Zongjian, et al.
Published: (2024)
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
by: Yuan, Shenghai, et al.
Published: (2025)
by: Yuan, Shenghai, et al.
Published: (2025)
ImgEdit: A Unified Image Editing Dataset and Benchmark
by: Ye, Yang, et al.
Published: (2025)
by: Ye, Yang, et al.
Published: (2025)
MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators
by: Yuan, Shenghai, et al.
Published: (2024)
by: Yuan, Shenghai, et al.
Published: (2024)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
by: Yuan, Shenghai, et al.
Published: (2024)
by: Yuan, Shenghai, et al.
Published: (2024)
Open-Sora Plan: Open-Source Large Video Generation Model
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
by: Li, Zongjian, et al.
Published: (2025)
by: Li, Zongjian, et al.
Published: (2025)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
by: Zhao, Min, et al.
Published: (2024)
by: Zhao, Min, et al.
Published: (2024)
EF-VI: Enhancing End-Frame Injection for Video Inbetweening
by: Chen, Liuhan, et al.
Published: (2025)
by: Chen, Liuhan, et al.
Published: (2025)
Mobius: Text to Seamless Looping Video Generation via Latent Shift
by: Bi, Xiuli, et al.
Published: (2025)
by: Bi, Xiuli, et al.
Published: (2025)
Conditional entropy for Amenable group actions
by: Lian, Yuan, et al.
Published: (2025)
by: Lian, Yuan, et al.
Published: (2025)
Latent Space Single-Pixel Imaging Under Low-Sampling Conditions
by: Yuan, Chenyu
Published: (2025)
by: Yuan, Chenyu
Published: (2025)
Breaking Latent Prior Bias in Detectors for Generalizable AIGC Image Detection
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
Jacquard V2: Refining Datasets using the Human In the Loop Data Correction Method
by: Li, Qiuhao, et al.
Published: (2024)
by: Li, Qiuhao, et al.
Published: (2024)
Are Conditional Latent Diffusion Models Effective for Image Restoration?
by: Yuan, Yunchen, et al.
Published: (2024)
by: Yuan, Yunchen, et al.
Published: (2024)
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
by: Yuan, Shenghai, et al.
Published: (2024)
by: Yuan, Shenghai, et al.
Published: (2024)
Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
by: Tang, Zhenyu, et al.
Published: (2024)
by: Tang, Zhenyu, et al.
Published: (2024)
Shifted Window Fourier Transform And Retention For Image Captioning
by: Hu, Jia Cheng, et al.
Published: (2024)
by: Hu, Jia Cheng, et al.
Published: (2024)
FlashVideo: A Framework for Swift Inference in Text-to-Video Generation
by: Lei, Bin, et al.
Published: (2023)
by: Lei, Bin, et al.
Published: (2023)
DLF: Extreme Image Compression with Dual-generative Latent Fusion
by: Xue, Naifu, et al.
Published: (2025)
by: Xue, Naifu, et al.
Published: (2025)
iFSQ: Improving FSQ for Image Generation with 1 Line of Code
by: Lin, Bin, et al.
Published: (2026)
by: Lin, Bin, et al.
Published: (2026)
RoomPainter: View-Integrated Diffusion for Consistent Indoor Scene Texturing
by: Huang, Zhipeng, et al.
Published: (2024)
by: Huang, Zhipeng, et al.
Published: (2024)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
Generative Latent Coding for Ultra-Low Bitrate Image and Video Compression
by: Qi, Linfeng, et al.
Published: (2025)
by: Qi, Linfeng, et al.
Published: (2025)
Progressive Image Restoration via Text-Conditioned Video Generation
by: Kang, Peng, et al.
Published: (2025)
by: Kang, Peng, et al.
Published: (2025)
MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
by: Zhang, Shuai, et al.
Published: (2025)
by: Zhang, Shuai, et al.
Published: (2025)
Watch Your Step: Learning Semantically-Guided Locomotion in Cluttered Environment
by: Liang, Denan, et al.
Published: (2026)
by: Liang, Denan, et al.
Published: (2026)
Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion Models
by: Zhu, Lin, et al.
Published: (2025)
by: Zhu, Lin, et al.
Published: (2025)
Latent Feature-Guided Conditional Diffusion for Generative Image Semantic Communication
by: Chen, Zehao, et al.
Published: (2025)
by: Chen, Zehao, et al.
Published: (2025)
Structured Task Solving via Modular Embodied Intelligence: A Case Study on Rubik's Cube
by: Fan, Chongshan, et al.
Published: (2025)
by: Fan, Chongshan, et al.
Published: (2025)
Biostatisticians Meet AI : Navigating Shifts While Preserving Principles
by: Bin Zhu
Published: (2025)
by: Bin Zhu
Published: (2025)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
by: Hu, Yaosi, et al.
Published: (2023)
by: Hu, Yaosi, et al.
Published: (2023)
ResVR: Joint Rescaling and Viewport Rendering of Omnidirectional Images
by: Li, Weiqi, et al.
Published: (2024)
by: Li, Weiqi, et al.
Published: (2024)
RPMS: Enhancing LLM-Based Embodied Planning through Rule-Augmented Memory Synergy
by: Yuan, Zhenhang, et al.
Published: (2026)
by: Yuan, Zhenhang, et al.
Published: (2026)
BEV-LIO(LC): BEV Image Assisted LiDAR-Inertial Odometry with Loop Closure
by: Cai, Haoxin, et al.
Published: (2025)
by: Cai, Haoxin, et al.
Published: (2025)
Similar Items
-
OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning
by: Ge, Yunyang, et al.
Published: (2026) -
SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video
by: Zhao, Chengshu, et al.
Published: (2025) -
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
by: Lin, Bin, et al.
Published: (2025) -
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
by: Chen, Liuhan, et al.
Published: (2024) -
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)