Relevance-guided Audio Visual Fusion for Video Saliency Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Li, Sun, Xuanzhe, Gao, Pan, Gabbouj, Moncef |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
by: Hooshanfar, Kiana, et al.
Published: (2025)
by: Hooshanfar, Kiana, et al.
Published: (2025)
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
FANeRV: Frequency Separation and Augmentation based Neural Representation for Video
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)
by: Cokelek, Mert, et al.
Published: (2025)
Degradation-Aware Hierarchical Termination for Blind Quality Enhancement of Compressed Video
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
Dropout Concrete Autoencoder for Band Selection on HSI Scenes
by: Xu, Lei, et al.
Published: (2024)
by: Xu, Lei, et al.
Published: (2024)
High-Frequency Enhanced Hybrid Neural Representation for Video Compression
by: Yu, Li, et al.
Published: (2024)
by: Yu, Li, et al.
Published: (2024)
Channel-wise Feature Decorrelation for Enhanced Learned Image Compression
by: Pakdaman, Farhad, et al.
Published: (2024)
by: Pakdaman, Farhad, et al.
Published: (2024)
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
by: Xiong, Junwen, et al.
Published: (2024)
by: Xiong, Junwen, et al.
Published: (2024)
Revisiting Generative Adversarial Networks for Binary Semantic Segmentation on Imbalanced Datasets
by: Xu, Lei, et al.
Published: (2024)
by: Xu, Lei, et al.
Published: (2024)
MMA-DFER: MultiModal Adaptation of unimodal models for Dynamic Facial Expression Recognition in-the-wild
by: Chumachenko, Kateryna, et al.
Published: (2024)
by: Chumachenko, Kateryna, et al.
Published: (2024)
Operational Support Estimator Networks
by: Ahishali, Mete, et al.
Published: (2023)
by: Ahishali, Mete, et al.
Published: (2023)
Multi-Scale Tensorial Summation and Dimensional Reduction Guided Neural Network for Edge Detection
by: Xu, Lei, et al.
Published: (2025)
by: Xu, Lei, et al.
Published: (2025)
Panoramic Image Inpainting With Gated Convolution And Contextual Reconstruction Loss
by: Yu, Li, et al.
Published: (2024)
by: Yu, Li, et al.
Published: (2024)
Efficient Audio-Visual Fusion for Video Classification
by: Awan, Mahrukh, et al.
Published: (2024)
by: Awan, Mahrukh, et al.
Published: (2024)
Joint End-to-End Image Compression and Denoising: Leveraging Contrastive Learning and Multi-Scale Self-ONNs
by: Xie, Yuxin, et al.
Published: (2024)
by: Xie, Yuxin, et al.
Published: (2024)
Perceptual Learned Image Compression via End-to-End JND-Based Optimization
by: Pakdaman, Farhad, et al.
Published: (2024)
by: Pakdaman, Farhad, et al.
Published: (2024)
KVQ: Boosting Video Quality Assessment via Saliency-guided Local Perception
by: Qu, Yunpeng, et al.
Published: (2025)
by: Qu, Yunpeng, et al.
Published: (2025)
Hyperspectral Image Analysis with Subspace Learning-based One-Class Classification
by: Kilickaya, Sertac, et al.
Published: (2023)
by: Kilickaya, Sertac, et al.
Published: (2023)
SRL-SOA: Self-Representation Learning with Sparse 1D-Operational Autoencoder for Hyperspectral Image Band Selection
by: Ahishali, Mete, et al.
Published: (2022)
by: Ahishali, Mete, et al.
Published: (2022)
Classification of Polarimetric SAR Images Using Compact Convolutional Neural Networks
by: Ahishali, Mete, et al.
Published: (2020)
by: Ahishali, Mete, et al.
Published: (2020)
Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
by: Gao, Yongbiao, et al.
Published: (2024)
by: Gao, Yongbiao, et al.
Published: (2024)
Crucial-Diff: A Unified Diffusion Model for Crucial Image and Annotation Synthesis in Data-scarce Scenarios
by: Yao, Siyue, et al.
Published: (2025)
by: Yao, Siyue, et al.
Published: (2025)
SAF-Net: Self-Attention Fusion Network for Myocardial Infarction Detection using Multi-View Echocardiography
by: Adalioglu, Ilke, et al.
Published: (2023)
by: Adalioglu, Ilke, et al.
Published: (2023)
Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification
by: Awan, Mahrukh, et al.
Published: (2024)
by: Awan, Mahrukh, et al.
Published: (2024)
Representation Based Regression for Object Distance Estimation
by: Ahishali, Mete, et al.
Published: (2021)
by: Ahishali, Mete, et al.
Published: (2021)
UL-DD: A Multimodal Drowsiness Dataset Using Video, Biometric Signals, and Behavioral Data
by: Bodaghi, Morteza, et al.
Published: (2025)
by: Bodaghi, Morteza, et al.
Published: (2025)
ViSAGE @ NTIRE 2026 Challenge on Video Saliency Prediction
by: Wang, Kun, et al.
Published: (2026)
by: Wang, Kun, et al.
Published: (2026)
Fusion-CAM: Integrating Gradient and Region-Based Class Activation Maps for Robust Visual Explanations
by: Dekdegue, Hajar, et al.
Published: (2026)
by: Dekdegue, Hajar, et al.
Published: (2026)
ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement
by: Ye, Jianping, et al.
Published: (2026)
by: Ye, Jianping, et al.
Published: (2026)
Contextual Encoder-Decoder Network for Visual Saliency Prediction
by: Kroner, Alexander, et al.
Published: (2019)
by: Kroner, Alexander, et al.
Published: (2019)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
by: Li, Jie, et al.
Published: (2023)
by: Li, Jie, et al.
Published: (2023)
Convolutional Sparse Support Estimator Network (CSEN) From energy efficient support estimation to learning-aided Compressive Sensing
by: Yamac, Mehmet, et al.
Published: (2020)
by: Yamac, Mehmet, et al.
Published: (2020)
Varying Manifolds in Diffusion: From Time-varying Geometries to Visual Saliency
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
Pixel-Wise Color Constancy via Smoothness Techniques in Multi-Illuminant Scenes
by: Entok, Umut Cem, et al.
Published: (2024)
by: Entok, Umut Cem, et al.
Published: (2024)
A Personalized Zero-Shot ECG Arrhythmia Monitoring System: From Sparse Representation Based Domain Adaption to Energy Efficient Abnormal Beat Detection for Practical ECG Surveillance
by: Yamaç, Mehmet, et al.
Published: (2022)
by: Yamaç, Mehmet, et al.
Published: (2022)
Saliency-guided Emotion Modeling: Predicting Viewer Reactions from Video Stimuli
by: Yaragoppa, Akhila, et al.
Published: (2025)
by: Yaragoppa, Akhila, et al.
Published: (2025)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
by: Ying, Xinru, et al.
Published: (2025)
by: Ying, Xinru, et al.
Published: (2025)
MTS-CSNet: Multiscale Tensor Factorization for Deep Compressive Sensing on RGB Images
by: Yamac, Mehmet, et al.
Published: (2026)
by: Yamac, Mehmet, et al.
Published: (2026)
Similar Items
-
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2025) -
DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
by: Hooshanfar, Kiana, et al.
Published: (2025) -
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
by: Yu, Li, et al.
Published: (2025) -
FANeRV: Frequency Separation and Augmentation based Neural Representation for Video
by: Yu, Li, et al.
Published: (2025) -
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)