RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Aiyue, Dong, Bin, Li, Jingru, Lin, Jing, Tian, Kun, Yao, Yiwu, Wang, Gongyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RainFusion2.0: Temporal-Spatial Awareness and Hardware-Efficient Block-wise Sparse Attention
by: Chen, Aiyue, et al.
Published: (2025)
by: Chen, Aiyue, et al.
Published: (2025)
Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers
by: Liu, Haosong, et al.
Published: (2025)
by: Liu, Haosong, et al.
Published: (2025)
NightRain: Nighttime Video Deraining via Adaptive-Rain-Removal and Adaptive-Correction
by: Lin, Beibei, et al.
Published: (2024)
by: Lin, Beibei, et al.
Published: (2024)
TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
by: Miao, Wenxuan, et al.
Published: (2025)
by: Miao, Wenxuan, et al.
Published: (2025)
Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
by: Li, Qirui, et al.
Published: (2025)
by: Li, Qirui, et al.
Published: (2025)
An Adaptive Underwater Image Enhancement Framework via Multi-Domain Fusion and Color Compensation
by: Tian, Yuezhe, et al.
Published: (2025)
by: Tian, Yuezhe, et al.
Published: (2025)
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
by: Liu, Yaoli, et al.
Published: (2025)
by: Liu, Yaoli, et al.
Published: (2025)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
by: Xing, Long, et al.
Published: (2024)
by: Xing, Long, et al.
Published: (2024)
ProvRain: Rain-Adaptive Denoising and Vehicle Detection via MobileNet-UNet and Faster R-CNN
by: Varathakumaran, Aswinkumar, et al.
Published: (2025)
by: Varathakumaran, Aswinkumar, et al.
Published: (2025)
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
by: Zhong, Yiwu, et al.
Published: (2024)
by: Zhong, Yiwu, et al.
Published: (2024)
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
by: Lee, Yuna, et al.
Published: (2026)
by: Lee, Yuna, et al.
Published: (2026)
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
by: Jian, Yiren, et al.
Published: (2023)
by: Jian, Yiren, et al.
Published: (2023)
AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction
by: Chen, Anjun, et al.
Published: (2024)
by: Chen, Anjun, et al.
Published: (2024)
Multi-granular body modeling with Redundancy-Free Spatiotemporal Fusion for Text-Driven Motion Generation
by: Zhan, Xingzu, et al.
Published: (2025)
by: Zhan, Xingzu, et al.
Published: (2025)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025)
by: Qiu, Haonan, et al.
Published: (2025)
Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
by: Qu, Chenfan, et al.
Published: (2025)
by: Qu, Chenfan, et al.
Published: (2025)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
by: Xiong, Tianwei, et al.
Published: (2026)
by: Xiong, Tianwei, et al.
Published: (2026)
Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
by: Dong, Haotian, et al.
Published: (2025)
by: Dong, Haotian, et al.
Published: (2025)
New Encoder Learning for Captioning Heavy Rain Images via Semantic Visual Feature Matching
by: Son, Chang-Hwan, et al.
Published: (2021)
by: Son, Chang-Hwan, et al.
Published: (2021)
Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
by: Cheng, Junhao, et al.
Published: (2026)
by: Cheng, Junhao, et al.
Published: (2026)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Learning Adaptive Fusion Bank for Multi-modal Salient Object Detection
by: Wang, Kunpeng, et al.
Published: (2024)
by: Wang, Kunpeng, et al.
Published: (2024)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
by: Ying, Xinru, et al.
Published: (2025)
by: Ying, Xinru, et al.
Published: (2025)
Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection
by: Lin, Kaiqing, et al.
Published: (2024)
by: Lin, Kaiqing, et al.
Published: (2024)
Text-Animator: Controllable Visual Text Video Generation
by: Liu, Lin, et al.
Published: (2024)
by: Liu, Lin, et al.
Published: (2024)
Harnessing Joint Rain-/Detail-aware Representations to Eliminate Intricate Rains
by: Ran, Wu, et al.
Published: (2024)
by: Ran, Wu, et al.
Published: (2024)
SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
by: Tong, Yujia, et al.
Published: (2026)
by: Tong, Yujia, et al.
Published: (2026)
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
by: Lu, Yu, et al.
Published: (2025)
by: Lu, Yu, et al.
Published: (2025)
HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention
by: Zheng, Xuzhe, et al.
Published: (2026)
by: Zheng, Xuzhe, et al.
Published: (2026)
Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion
by: Wang, Mengyu, et al.
Published: (2025)
by: Wang, Mengyu, et al.
Published: (2025)
Efficient Audio-Visual Fusion for Video Classification
by: Awan, Mahrukh, et al.
Published: (2024)
by: Awan, Mahrukh, et al.
Published: (2024)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
by: Li, Hongliang, et al.
Published: (2025)
by: Li, Hongliang, et al.
Published: (2025)
L2SR: Learning to Sample and Reconstruct for Accelerated MRI via Reinforcement Learning
by: Yang, Pu, et al.
Published: (2022)
by: Yang, Pu, et al.
Published: (2022)
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
by: An, Zhaochong, et al.
Published: (2025)
by: An, Zhaochong, et al.
Published: (2025)
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
by: Shi, Jiaqi, et al.
Published: (2026)
by: Shi, Jiaqi, et al.
Published: (2026)
Revisiting Tampered Scene Text Detection in the Era of Generative AI
by: Qu, Chenfan, et al.
Published: (2024)
by: Qu, Chenfan, et al.
Published: (2024)
RainMamba: Enhanced Locality Learning with State Space Models for Video Deraining
by: Wu, Hongtao, et al.
Published: (2024)
by: Wu, Hongtao, et al.
Published: (2024)
Similar Items
-
RainFusion2.0: Temporal-Spatial Awareness and Hardware-Efficient Block-wise Sparse Attention
by: Chen, Aiyue, et al.
Published: (2025) -
Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers
by: Liu, Haosong, et al.
Published: (2025) -
NightRain: Nighttime Video Deraining via Adaptive-Rain-Removal and Adaptive-Correction
by: Lin, Beibei, et al.
Published: (2024) -
TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
by: Miao, Wenxuan, et al.
Published: (2025) -
Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
by: Li, Qirui, et al.
Published: (2025)