VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Linfeng, Wang, Yeda, Gong, Meiqi, Li, Zizhuo, Deng, Yuxin, Yi, Xunpeng, Li, Chunyu, Xu, Han, Zhang, Hao, Ma, Jiayi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
by: Gong, Meiqi, et al.
Published: (2025)
by: Gong, Meiqi, et al.
Published: (2025)
Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
by: Yi, Xunpeng, et al.
Published: (2024)
by: Yi, Xunpeng, et al.
Published: (2024)
MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
ControlFusion: A Controllable Image Fusion Framework with Language-Vision Degradation Prompts
by: Tang, Linfeng, et al.
Published: (2025)
by: Tang, Linfeng, et al.
Published: (2025)
Efficient Sparse-to-Dense Visual Localization via Compact Gaussian Scene Representation and Accelerated Dense Pose Estimation
by: Li, Zizhuo, et al.
Published: (2026)
by: Li, Zizhuo, et al.
Published: (2026)
DSPFusion: Image Fusion via Degradation and Semantic Dual-Prior Guidance
by: Tang, Linfeng, et al.
Published: (2025)
by: Tang, Linfeng, et al.
Published: (2025)
LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up Tables
by: Yi, Xunpeng, et al.
Published: (2025)
by: Yi, Xunpeng, et al.
Published: (2025)
STF: Spatio-Temporal Fusion Module for Improving Video Object Detection
by: Anwar, Noreen, et al.
Published: (2024)
by: Anwar, Noreen, et al.
Published: (2024)
STAF: 3D Human Mesh Recovery from Video with Spatio-Temporal Alignment Fusion
by: Yao, Wei, et al.
Published: (2024)
by: Yao, Wei, et al.
Published: (2024)
CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image Matching
by: Li, Zizhuo, et al.
Published: (2025)
by: Li, Zizhuo, et al.
Published: (2025)
Deep Learning Reforms Image Matching: A Survey and Outlook
by: Zhang, Shihua, et al.
Published: (2025)
by: Zhang, Shihua, et al.
Published: (2025)
Fusion of Spatio-Temporal and Multi-Scale Frequency Features for Dry Electrodes MI-EEG Decoding
by: Gong, Tianyi, et al.
Published: (2026)
by: Gong, Tianyi, et al.
Published: (2026)
Global Spatio-Temporal Fusion-based Traffic Prediction Algorithm with Anomaly Aware
by: Liu, Chaoqun, et al.
Published: (2024)
by: Liu, Chaoqun, et al.
Published: (2024)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
by: Ying, Xinru, et al.
Published: (2025)
by: Ying, Xinru, et al.
Published: (2025)
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
MapFusion: A Novel BEV Feature Fusion Network for Multi-modal Map Construction
by: Hao, Xiaoshuai, et al.
Published: (2025)
by: Hao, Xiaoshuai, et al.
Published: (2025)
DCG ReID: Disentangling Collaboration and Guidance Fusion Representations for Multi-modal Vehicle Re-Identification
by: Zheng, Aihua, et al.
Published: (2026)
by: Zheng, Aihua, et al.
Published: (2026)
MCSFF: Multi-modal Consistency and Specificity Fusion Framework for Entity Alignment
by: Ai, Wei, et al.
Published: (2024)
by: Ai, Wei, et al.
Published: (2024)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
by: Cheng, Zixu, et al.
Published: (2025)
by: Cheng, Zixu, et al.
Published: (2025)
A Multi-modal Fusion Network for Terrain Perception Based on Illumination Aware
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
FTPFusion: Frequency-Aware Infrared and Visible Video Fusion with Temporal Perturbation
by: Li, Xilai, et al.
Published: (2026)
by: Li, Xilai, et al.
Published: (2026)
Spiking Neural Networks with Temporal Attention-Guided Adaptive Fusion for imbalanced Multi-modal Learning
by: Shen, Jiangrong, et al.
Published: (2025)
by: Shen, Jiangrong, et al.
Published: (2025)
ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion
by: Zhu, Jun, et al.
Published: (2025)
by: Zhu, Jun, et al.
Published: (2025)
Multi-modal Fusion based Q-distribution Prediction for Controlled Nuclear Fusion
by: Wang, Shiao, et al.
Published: (2024)
by: Wang, Shiao, et al.
Published: (2024)
Spatial Decomposition and Temporal Fusion based Inter Prediction for Learned Video Compression
by: Sheng, Xihua, et al.
Published: (2024)
by: Sheng, Xihua, et al.
Published: (2024)
Dissecting RGB-D Learning for Improved Multi-modal Fusion
by: Chen, Hao, et al.
Published: (2023)
by: Chen, Hao, et al.
Published: (2023)
DRFusion: Drift-Resilient Temporally Consistent Infrared-Visible Video Fusion
by: Li, Xingyuan, et al.
Published: (2026)
by: Li, Xingyuan, et al.
Published: (2026)
SpatioTemporal Difference Network for Video Depth Super-Resolution
by: Wang, Zhengxue, et al.
Published: (2025)
by: Wang, Zhengxue, et al.
Published: (2025)
Student Classroom Behavior Detection based on Spatio-Temporal Network and Multi-Model Fusion
by: Yang, Fan, et al.
Published: (2023)
by: Yang, Fan, et al.
Published: (2023)
Collaborative Cross-modal Fusion with Large Language Model for Recommendation
by: Liu, Zhongzhou, et al.
Published: (2024)
by: Liu, Zhongzhou, et al.
Published: (2024)
Standardization for improved Spatio-Temporal Image Fusion
by: Goyena, Harkaitz, et al.
Published: (2025)
by: Goyena, Harkaitz, et al.
Published: (2025)
Co-Fusion4D: Spatio-temporal Collaborative Fusion for Robust 3D Object Detection
by: Li, Wenxuan, et al.
Published: (2026)
by: Li, Wenxuan, et al.
Published: (2026)
Apollo: Unified Multi-Task Audio-Video Joint Generation
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
Integrated CNN‐LSTM for Photovoltaic Power Prediction based on Spatio‐Temporal Feature Fusion
by: Junwei Ma, et al.
Published: (2024)
by: Junwei Ma, et al.
Published: (2024)
Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper
by: Li, Ling, et al.
Published: (2026)
by: Li, Ling, et al.
Published: (2026)
3D Multi-frame Fusion for Video Stabilization
by: Peng, Zhan, et al.
Published: (2024)
by: Peng, Zhan, et al.
Published: (2024)
Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search
by: Tang, Hengzhu, et al.
Published: (2025)
by: Tang, Hengzhu, et al.
Published: (2025)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
by: Zhang, Junbin, et al.
Published: (2022)
by: Zhang, Junbin, et al.
Published: (2022)
Similar Items
-
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
by: Gong, Meiqi, et al.
Published: (2025) -
Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
by: Yi, Xunpeng, et al.
Published: (2024) -
MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement
by: Zhang, Hao, et al.
Published: (2026) -
ControlFusion: A Controllable Image Fusion Framework with Language-Vision Degradation Prompts
by: Tang, Linfeng, et al.
Published: (2025) -
Efficient Sparse-to-Dense Visual Localization via Compact Gaussian Scene Representation and Accelerated Dense Pose Estimation
by: Li, Zizhuo, et al.
Published: (2026)