DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hooshanfar, Kiana, Hosseini, Alireza, Kalhor, Ahmad, Araabi, Babak Nadjar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Patch-Wise Self-Supervised Visual Representation Learning: A Fine-Grained Approach
von: Javidani, Ali, et al.
Veröffentlicht: (2023)
von: Javidani, Ali, et al.
Veröffentlicht: (2023)
Brand Visibility in Packaging: A Deep Learning Approach for Logo Detection, Saliency-Map Prediction, and Logo Placement Analysis
von: Hosseini, Alireza, et al.
Veröffentlicht: (2024)
von: Hosseini, Alireza, et al.
Veröffentlicht: (2024)
Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning
von: Javidani, Ali, et al.
Veröffentlicht: (2025)
von: Javidani, Ali, et al.
Veröffentlicht: (2025)
Relevance-guided Audio Visual Fusion for Video Saliency Prediction
von: Yu, Li, et al.
Veröffentlicht: (2024)
von: Yu, Li, et al.
Veröffentlicht: (2024)
Robustifying Diffusion-Denoised Smoothing Against Covariate Shift
von: Hedayatnia, Ali, et al.
Veröffentlicht: (2025)
von: Hedayatnia, Ali, et al.
Veröffentlicht: (2025)
SUM: Saliency Unification through Mamba for Visual Attention Modeling
von: Hosseini, Alireza, et al.
Veröffentlicht: (2024)
von: Hosseini, Alireza, et al.
Veröffentlicht: (2024)
Enhancing Interpretability of Sparse Latent Representations with Class Information
von: Abiz, Farshad Sangari, et al.
Veröffentlicht: (2025)
von: Abiz, Farshad Sangari, et al.
Veröffentlicht: (2025)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
von: Yu, Li, et al.
Veröffentlicht: (2025)
von: Yu, Li, et al.
Veröffentlicht: (2025)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
von: Cokelek, Mert, et al.
Veröffentlicht: (2025)
von: Cokelek, Mert, et al.
Veröffentlicht: (2025)
Improving the Successful Robotic Grasp Detection Using Convolutional Neural Networks
von: Hosseini, Hamed, et al.
Veröffentlicht: (2024)
von: Hosseini, Hamed, et al.
Veröffentlicht: (2024)
AI-Driven Relocation Tracking in Dynamic Kitchen Environments
von: Esfahani, Arash Nasr, et al.
Veröffentlicht: (2025)
von: Esfahani, Arash Nasr, et al.
Veröffentlicht: (2025)
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
von: Xiong, Junwen, et al.
Veröffentlicht: (2024)
von: Xiong, Junwen, et al.
Veröffentlicht: (2024)
Efficient Audio-Visual Fusion for Video Classification
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement
von: Ye, Jianping, et al.
Veröffentlicht: (2026)
von: Ye, Jianping, et al.
Veröffentlicht: (2026)
SalFoM: Dynamic Saliency Prediction with Video Foundation Models
von: Moradi, Morteza, et al.
Veröffentlicht: (2024)
von: Moradi, Morteza, et al.
Veröffentlicht: (2024)
Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
Contextual Encoder-Decoder Network for Visual Saliency Prediction
von: Kroner, Alexander, et al.
Veröffentlicht: (2019)
von: Kroner, Alexander, et al.
Veröffentlicht: (2019)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
Salience-Based Adaptive Masking: Revisiting Token Dynamics for Enhanced Pre-training
von: Choi, Hyesong, et al.
Veröffentlicht: (2024)
von: Choi, Hyesong, et al.
Veröffentlicht: (2024)
ViSAGE @ NTIRE 2026 Challenge on Video Saliency Prediction
von: Wang, Kun, et al.
Veröffentlicht: (2026)
von: Wang, Kun, et al.
Veröffentlicht: (2026)
Scene Understanding in Pick-and-Place Tasks: Analyzing Transformations Between Initial and Final Scenes
von: Ghasemi, Seraj, et al.
Veröffentlicht: (2024)
von: Ghasemi, Seraj, et al.
Veröffentlicht: (2024)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
von: Choi, Seung hee, et al.
Veröffentlicht: (2026)
von: Choi, Seung hee, et al.
Veröffentlicht: (2026)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
von: Li, Jie, et al.
Veröffentlicht: (2023)
von: Li, Jie, et al.
Veröffentlicht: (2023)
Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
von: Yu, Yinfeng, et al.
Veröffentlicht: (2025)
von: Yu, Yinfeng, et al.
Veröffentlicht: (2025)
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition
von: Cong, Kaixuan, et al.
Veröffentlicht: (2025)
von: Cong, Kaixuan, et al.
Veröffentlicht: (2025)
AVT2-DWF: Improving Deepfake Detection with Audio-Visual Fusion and Dynamic Weighting Strategies
von: Wang, Rui, et al.
Veröffentlicht: (2024)
von: Wang, Rui, et al.
Veröffentlicht: (2024)
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
von: Moradi, Morteza, et al.
Veröffentlicht: (2024)
von: Moradi, Morteza, et al.
Veröffentlicht: (2024)
Minimalistic Video Saliency Prediction via Efficient Decoder & Spatio Temporal Action Cues
von: Girmaji, Rohit, et al.
Veröffentlicht: (2025)
von: Girmaji, Rohit, et al.
Veröffentlicht: (2025)
Enhancing Saliency Prediction in Monitoring Tasks: The Role of Visual Highlights
von: Wu, Zekun, et al.
Veröffentlicht: (2024)
von: Wu, Zekun, et al.
Veröffentlicht: (2024)
SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
von: Deng, Jinhong, et al.
Veröffentlicht: (2025)
von: Deng, Jinhong, et al.
Veröffentlicht: (2025)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
von: Jin, Dian, et al.
Veröffentlicht: (2025)
von: Jin, Dian, et al.
Veröffentlicht: (2025)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
von: Wang, Junke, et al.
Veröffentlicht: (2024)
von: Wang, Junke, et al.
Veröffentlicht: (2024)
Leveraging Multi-Modal Saliency and Fusion for Gaze Target Detection
von: Mathew, Athul M., et al.
Veröffentlicht: (2025)
von: Mathew, Athul M., et al.
Veröffentlicht: (2025)
VST++: Efficient and Stronger Visual Saliency Transformer
von: Liu, Nian, et al.
Veröffentlicht: (2023)
von: Liu, Nian, et al.
Veröffentlicht: (2023)
Finding Visual Saliency in Continuous Spike Stream
von: Zhu, Lin, et al.
Veröffentlicht: (2024)
von: Zhu, Lin, et al.
Veröffentlicht: (2024)
Principles of Visual Tokens for Efficient Video Understanding
von: Hao, Xinyue, et al.
Veröffentlicht: (2024)
von: Hao, Xinyue, et al.
Veröffentlicht: (2024)
CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2024)
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Patch-Wise Self-Supervised Visual Representation Learning: A Fine-Grained Approach
von: Javidani, Ali, et al.
Veröffentlicht: (2023) -
Brand Visibility in Packaging: A Deep Learning Approach for Logo Detection, Saliency-Map Prediction, and Logo Placement Analysis
von: Hosseini, Alireza, et al.
Veröffentlicht: (2024) -
Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning
von: Javidani, Ali, et al.
Veröffentlicht: (2025) -
Relevance-guided Audio Visual Fusion for Video Saliency Prediction
von: Yu, Li, et al.
Veröffentlicht: (2024) -
Robustifying Diffusion-Denoised Smoothing Against Covariate Shift
von: Hedayatnia, Ali, et al.
Veröffentlicht: (2025)