Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Li, Sun, Xuanzhe, Zhou, Wei, Gabbouj, Moncef |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Relevance-guided Audio Visual Fusion for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2024)
by: Yu, Li, et al.
Published: (2024)
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
by: Xiong, Junwen, et al.
Published: (2024)
by: Xiong, Junwen, et al.
Published: (2024)
FANeRV: Frequency Separation and Augmentation based Neural Representation for Video
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
by: Hooshanfar, Kiana, et al.
Published: (2025)
by: Hooshanfar, Kiana, et al.
Published: (2025)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)
by: Cokelek, Mert, et al.
Published: (2025)
Degradation-Aware Hierarchical Termination for Blind Quality Enhancement of Compressed Video
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
Dropout Concrete Autoencoder for Band Selection on HSI Scenes
by: Xu, Lei, et al.
Published: (2024)
by: Xu, Lei, et al.
Published: (2024)
High-Frequency Enhanced Hybrid Neural Representation for Video Compression
by: Yu, Li, et al.
Published: (2024)
by: Yu, Li, et al.
Published: (2024)
Crucial-Diff: A Unified Diffusion Model for Crucial Image and Annotation Synthesis in Data-scarce Scenarios
by: Yao, Siyue, et al.
Published: (2025)
by: Yao, Siyue, et al.
Published: (2025)
Channel-wise Feature Decorrelation for Enhanced Learned Image Compression
by: Pakdaman, Farhad, et al.
Published: (2024)
by: Pakdaman, Farhad, et al.
Published: (2024)
Revisiting Generative Adversarial Networks for Binary Semantic Segmentation on Imbalanced Datasets
by: Xu, Lei, et al.
Published: (2024)
by: Xu, Lei, et al.
Published: (2024)
MMA-DFER: MultiModal Adaptation of unimodal models for Dynamic Facial Expression Recognition in-the-wild
by: Chumachenko, Kateryna, et al.
Published: (2024)
by: Chumachenko, Kateryna, et al.
Published: (2024)
Multi-Scale Tensorial Summation and Dimensional Reduction Guided Neural Network for Edge Detection
by: Xu, Lei, et al.
Published: (2025)
by: Xu, Lei, et al.
Published: (2025)
Operational Support Estimator Networks
by: Ahishali, Mete, et al.
Published: (2023)
by: Ahishali, Mete, et al.
Published: (2023)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
by: Sun, Zhengyang, et al.
Published: (2026)
by: Sun, Zhengyang, et al.
Published: (2026)
Joint End-to-End Image Compression and Denoising: Leveraging Contrastive Learning and Multi-Scale Self-ONNs
by: Xie, Yuxin, et al.
Published: (2024)
by: Xie, Yuxin, et al.
Published: (2024)
Perceptual Learned Image Compression via End-to-End JND-Based Optimization
by: Pakdaman, Farhad, et al.
Published: (2024)
by: Pakdaman, Farhad, et al.
Published: (2024)
Hyperspectral Image Analysis with Subspace Learning-based One-Class Classification
by: Kilickaya, Sertac, et al.
Published: (2023)
by: Kilickaya, Sertac, et al.
Published: (2023)
SRL-SOA: Self-Representation Learning with Sparse 1D-Operational Autoencoder for Hyperspectral Image Band Selection
by: Ahishali, Mete, et al.
Published: (2022)
by: Ahishali, Mete, et al.
Published: (2022)
Classification of Polarimetric SAR Images Using Compact Convolutional Neural Networks
by: Ahishali, Mete, et al.
Published: (2020)
by: Ahishali, Mete, et al.
Published: (2020)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
by: Li, Jie, et al.
Published: (2023)
by: Li, Jie, et al.
Published: (2023)
CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
Saliency Guided Optimization of Diffusion Latents
by: Wang, Xiwen, et al.
Published: (2024)
by: Wang, Xiwen, et al.
Published: (2024)
Panoramic Image Inpainting With Gated Convolution And Contextual Reconstruction Loss
by: Yu, Li, et al.
Published: (2024)
by: Yu, Li, et al.
Published: (2024)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
by: Wang, Xingrui, et al.
Published: (2025)
by: Wang, Xingrui, et al.
Published: (2025)
Representation Based Regression for Object Distance Estimation
by: Ahishali, Mete, et al.
Published: (2021)
by: Ahishali, Mete, et al.
Published: (2021)
UL-DD: A Multimodal Drowsiness Dataset Using Video, Biometric Signals, and Behavioral Data
by: Bodaghi, Morteza, et al.
Published: (2025)
by: Bodaghi, Morteza, et al.
Published: (2025)
SalFoM: Dynamic Saliency Prediction with Video Foundation Models
by: Moradi, Morteza, et al.
Published: (2024)
by: Moradi, Morteza, et al.
Published: (2024)
ViSAGE @ NTIRE 2026 Challenge on Video Saliency Prediction
by: Wang, Kun, et al.
Published: (2026)
by: Wang, Kun, et al.
Published: (2026)
GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
by: Dou, Huanzhang, et al.
Published: (2024)
by: Dou, Huanzhang, et al.
Published: (2024)
ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement
by: Ye, Jianping, et al.
Published: (2026)
by: Ye, Jianping, et al.
Published: (2026)
Contextual Encoder-Decoder Network for Visual Saliency Prediction
by: Kroner, Alexander, et al.
Published: (2019)
by: Kroner, Alexander, et al.
Published: (2019)
Data Augmentation via Latent Diffusion for Saliency Prediction
by: Aydemir, Bahar, et al.
Published: (2024)
by: Aydemir, Bahar, et al.
Published: (2024)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
by: Sun, Yasheng, et al.
Published: (2025)
by: Sun, Yasheng, et al.
Published: (2025)
Convolutional Sparse Support Estimator Network (CSEN) From energy efficient support estimation to learning-aided Compressive Sensing
by: Yamac, Mehmet, et al.
Published: (2020)
by: Yamac, Mehmet, et al.
Published: (2020)
A Psychophysically Oriented Saliency Map Prediction Model
by: Li, Qiang
Published: (2020)
by: Li, Qiang
Published: (2020)
A Personalized Zero-Shot ECG Arrhythmia Monitoring System: From Sparse Representation Based Domain Adaption to Energy Efficient Abnormal Beat Detection for Practical ECG Surveillance
by: Yamaç, Mehmet, et al.
Published: (2022)
by: Yamaç, Mehmet, et al.
Published: (2022)
Pixel-Wise Color Constancy via Smoothness Techniques in Multi-Illuminant Scenes
by: Entok, Umut Cem, et al.
Published: (2024)
by: Entok, Umut Cem, et al.
Published: (2024)
Similar Items
-
Relevance-guided Audio Visual Fusion for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2024) -
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
by: Yu, Li, et al.
Published: (2025) -
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
by: Xiong, Junwen, et al.
Published: (2024) -
FANeRV: Frequency Separation and Augmentation based Neural Representation for Video
by: Yu, Li, et al.
Published: (2025) -
DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
by: Hooshanfar, Kiana, et al.
Published: (2025)