RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Aiyue, Dong, Bin, Li, Jingru, Lin, Jing, Tian, Kun, Yao, Yiwu, Wang, Gongyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RainFusion2.0: Temporal-Spatial Awareness and Hardware-Efficient Block-wise Sparse Attention
von: Chen, Aiyue, et al.
Veröffentlicht: (2025)
von: Chen, Aiyue, et al.
Veröffentlicht: (2025)
Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers
von: Liu, Haosong, et al.
Veröffentlicht: (2025)
von: Liu, Haosong, et al.
Veröffentlicht: (2025)
NightRain: Nighttime Video Deraining via Adaptive-Rain-Removal and Adaptive-Correction
von: Lin, Beibei, et al.
Veröffentlicht: (2024)
von: Lin, Beibei, et al.
Veröffentlicht: (2024)
TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
von: Miao, Wenxuan, et al.
Veröffentlicht: (2025)
von: Miao, Wenxuan, et al.
Veröffentlicht: (2025)
Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
von: Li, Qirui, et al.
Veröffentlicht: (2025)
von: Li, Qirui, et al.
Veröffentlicht: (2025)
An Adaptive Underwater Image Enhancement Framework via Multi-Domain Fusion and Color Compensation
von: Tian, Yuezhe, et al.
Veröffentlicht: (2025)
von: Tian, Yuezhe, et al.
Veröffentlicht: (2025)
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
von: Liu, Yaoli, et al.
Veröffentlicht: (2025)
von: Liu, Yaoli, et al.
Veröffentlicht: (2025)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
von: Xing, Long, et al.
Veröffentlicht: (2024)
von: Xing, Long, et al.
Veröffentlicht: (2024)
ProvRain: Rain-Adaptive Denoising and Vehicle Detection via MobileNet-UNet and Faster R-CNN
von: Varathakumaran, Aswinkumar, et al.
Veröffentlicht: (2025)
von: Varathakumaran, Aswinkumar, et al.
Veröffentlicht: (2025)
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
von: Lee, Yuna, et al.
Veröffentlicht: (2026)
von: Lee, Yuna, et al.
Veröffentlicht: (2026)
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
von: Jian, Yiren, et al.
Veröffentlicht: (2023)
von: Jian, Yiren, et al.
Veröffentlicht: (2023)
AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction
von: Chen, Anjun, et al.
Veröffentlicht: (2024)
von: Chen, Anjun, et al.
Veröffentlicht: (2024)
Multi-granular body modeling with Redundancy-Free Spatiotemporal Fusion for Text-Driven Motion Generation
von: Zhan, Xingzu, et al.
Veröffentlicht: (2025)
von: Zhan, Xingzu, et al.
Veröffentlicht: (2025)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
von: Qiu, Haonan, et al.
Veröffentlicht: (2025)
von: Qiu, Haonan, et al.
Veröffentlicht: (2025)
Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
von: Qu, Chenfan, et al.
Veröffentlicht: (2025)
von: Qu, Chenfan, et al.
Veröffentlicht: (2025)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
von: Xiong, Tianwei, et al.
Veröffentlicht: (2026)
von: Xiong, Tianwei, et al.
Veröffentlicht: (2026)
Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
New Encoder Learning for Captioning Heavy Rain Images via Semantic Visual Feature Matching
von: Son, Chang-Hwan, et al.
Veröffentlicht: (2021)
von: Son, Chang-Hwan, et al.
Veröffentlicht: (2021)
Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
von: Cheng, Junhao, et al.
Veröffentlicht: (2026)
von: Cheng, Junhao, et al.
Veröffentlicht: (2026)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
von: Wang, Kai, et al.
Veröffentlicht: (2024)
von: Wang, Kai, et al.
Veröffentlicht: (2024)
Learning Adaptive Fusion Bank for Multi-modal Salient Object Detection
von: Wang, Kunpeng, et al.
Veröffentlicht: (2024)
von: Wang, Kunpeng, et al.
Veröffentlicht: (2024)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
von: Ying, Xinru, et al.
Veröffentlicht: (2025)
von: Ying, Xinru, et al.
Veröffentlicht: (2025)
Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection
von: Lin, Kaiqing, et al.
Veröffentlicht: (2024)
von: Lin, Kaiqing, et al.
Veröffentlicht: (2024)
Text-Animator: Controllable Visual Text Video Generation
von: Liu, Lin, et al.
Veröffentlicht: (2024)
von: Liu, Lin, et al.
Veröffentlicht: (2024)
Harnessing Joint Rain-/Detail-aware Representations to Eliminate Intricate Rains
von: Ran, Wu, et al.
Veröffentlicht: (2024)
von: Ran, Wu, et al.
Veröffentlicht: (2024)
SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
von: Tong, Yujia, et al.
Veröffentlicht: (2026)
von: Tong, Yujia, et al.
Veröffentlicht: (2026)
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
von: Lu, Yu, et al.
Veröffentlicht: (2025)
von: Lu, Yu, et al.
Veröffentlicht: (2025)
HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention
von: Zheng, Xuzhe, et al.
Veröffentlicht: (2026)
von: Zheng, Xuzhe, et al.
Veröffentlicht: (2026)
Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion
von: Wang, Mengyu, et al.
Veröffentlicht: (2025)
von: Wang, Mengyu, et al.
Veröffentlicht: (2025)
Efficient Audio-Visual Fusion for Video Classification
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
von: Li, Hongliang, et al.
Veröffentlicht: (2025)
von: Li, Hongliang, et al.
Veröffentlicht: (2025)
L2SR: Learning to Sample and Reconstruct for Accelerated MRI via Reinforcement Learning
von: Yang, Pu, et al.
Veröffentlicht: (2022)
von: Yang, Pu, et al.
Veröffentlicht: (2022)
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
Revisiting Tampered Scene Text Detection in the Era of Generative AI
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
RainMamba: Enhanced Locality Learning with State Space Models for Video Deraining
von: Wu, Hongtao, et al.
Veröffentlicht: (2024)
von: Wu, Hongtao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RainFusion2.0: Temporal-Spatial Awareness and Hardware-Efficient Block-wise Sparse Attention
von: Chen, Aiyue, et al.
Veröffentlicht: (2025) -
Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers
von: Liu, Haosong, et al.
Veröffentlicht: (2025) -
NightRain: Nighttime Video Deraining via Adaptive-Rain-Removal and Adaptive-Correction
von: Lin, Beibei, et al.
Veröffentlicht: (2024) -
TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
von: Miao, Wenxuan, et al.
Veröffentlicht: (2025) -
Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
von: Li, Qirui, et al.
Veröffentlicht: (2025)