Minimalistic Video Saliency Prediction via Efficient Decoder & Spatio Temporal Action Cues
Fuente:
arXiv
Guardado en:
| Autores principales: | Girmaji, Rohit, Jain, Siddharth, Beri, Bhav, Bansal, Sarthak, Gandhi, Vineet |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EditIQ: Automated Cinematic Editing of Static Wide-Angle Videos via Dialogue Interpretation and Saliency Cues
por: Girmaji, Rohit, et al.
Publicado: (2025)
por: Girmaji, Rohit, et al.
Publicado: (2025)
Simplifying Knowledge Transfer in Pretrained Models
por: Jain, Siddharth, et al.
Publicado: (2025)
por: Jain, Siddharth, et al.
Publicado: (2025)
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
por: Moradi, Morteza, et al.
Publicado: (2024)
por: Moradi, Morteza, et al.
Publicado: (2024)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
por: Wang, Yubin, et al.
Publicado: (2024)
por: Wang, Yubin, et al.
Publicado: (2024)
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
por: Siddiqui, Nyle, et al.
Publicado: (2025)
por: Siddiqui, Nyle, et al.
Publicado: (2025)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
por: Wang, Xiao, et al.
Publicado: (2026)
por: Wang, Xiao, et al.
Publicado: (2026)
Saliency-guided Emotion Modeling: Predicting Viewer Reactions from Video Stimuli
por: Yaragoppa, Akhila, et al.
Publicado: (2025)
por: Yaragoppa, Akhila, et al.
Publicado: (2025)
Vid-Freeze: Protecting Images from Malicious Image-to-Video Generation via Temporal Freezing
por: Chowdhury, Rohit, et al.
Publicado: (2025)
por: Chowdhury, Rohit, et al.
Publicado: (2025)
Contextual Encoder-Decoder Network for Visual Saliency Prediction
por: Kroner, Alexander, et al.
Publicado: (2019)
por: Kroner, Alexander, et al.
Publicado: (2019)
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
por: Shen, Hao, et al.
Publicado: (2024)
por: Shen, Hao, et al.
Publicado: (2024)
Open-Vocabulary Spatio-Temporal Action Detection
por: Wu, Tao, et al.
Publicado: (2024)
por: Wu, Tao, et al.
Publicado: (2024)
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
por: Zhuang, Weijun, et al.
Publicado: (2026)
por: Zhuang, Weijun, et al.
Publicado: (2026)
Video-Language Alignment via Spatio-Temporal Graph Transformer
por: Zhang, Shi-Xue, et al.
Publicado: (2024)
por: Zhang, Shi-Xue, et al.
Publicado: (2024)
ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement
por: Ye, Jianping, et al.
Publicado: (2026)
por: Ye, Jianping, et al.
Publicado: (2026)
Beyond Pixels: Leveraging the Language of Soccer to Improve Spatio-Temporal Action Detection in Broadcast Videos
por: Ochin, Jeremie, et al.
Publicado: (2025)
por: Ochin, Jeremie, et al.
Publicado: (2025)
TempSAL -- Uncovering Temporal Information for Deep Saliency Prediction
por: Aydemir, Bahar, et al.
Publicado: (2023)
por: Aydemir, Bahar, et al.
Publicado: (2023)
Improving Temporal Action Segmentation via Constraint-Aware Decoding
por: Ee, Yeo Keat, et al.
Publicado: (2026)
por: Ee, Yeo Keat, et al.
Publicado: (2026)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
por: Rajendiran, Ramanathan, et al.
Publicado: (2023)
por: Rajendiran, Ramanathan, et al.
Publicado: (2023)
Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
por: Agarwal, Aishwarya, et al.
Publicado: (2025)
por: Agarwal, Aishwarya, et al.
Publicado: (2025)
TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time Correction
por: Agarwal, Aishwarya, et al.
Publicado: (2024)
por: Agarwal, Aishwarya, et al.
Publicado: (2024)
LiteEmbed: Adapting CLIP to Rare Classes
por: Agarwal, Aishwarya, et al.
Publicado: (2026)
por: Agarwal, Aishwarya, et al.
Publicado: (2026)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
por: Gao, Shida, et al.
Publicado: (2025)
por: Gao, Shida, et al.
Publicado: (2025)
PEEKABOO: Interactive Video Generation via Masked-Diffusion
por: Jain, Yash, et al.
Publicado: (2023)
por: Jain, Yash, et al.
Publicado: (2023)
Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
por: Zhuge, Yunzhi, et al.
Publicado: (2025)
por: Zhuge, Yunzhi, et al.
Publicado: (2025)
DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition
por: Ullah, Hayat, et al.
Publicado: (2025)
por: Ullah, Hayat, et al.
Publicado: (2025)
UniSTFormer: Unified Spatio-Temporal Lightweight Transformer for Efficient Skeleton-Based Action Recognition
por: Wu, Wenhan, et al.
Publicado: (2025)
por: Wu, Wenhan, et al.
Publicado: (2025)
Context-Guided Spatio-Temporal Video Grounding
por: Gu, Xin, et al.
Publicado: (2024)
por: Gu, Xin, et al.
Publicado: (2024)
ViSAGE @ NTIRE 2026 Challenge on Video Saliency Prediction
por: Wang, Kun, et al.
Publicado: (2026)
por: Wang, Kun, et al.
Publicado: (2026)
Relevance-guided Audio Visual Fusion for Video Saliency Prediction
por: Yu, Li, et al.
Publicado: (2024)
por: Yu, Li, et al.
Publicado: (2024)
V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
por: Lin, Xinying, et al.
Publicado: (2026)
por: Lin, Xinying, et al.
Publicado: (2026)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
por: Li, Peiyan, et al.
Publicado: (2026)
por: Li, Peiyan, et al.
Publicado: (2026)
Resolving Spatio-Temporal Entanglement in Video Prediction via Multi-Modal Attention
por: Gupta, Shreyam, et al.
Publicado: (2025)
por: Gupta, Shreyam, et al.
Publicado: (2025)
GS-STVSR: Ultra-Efficient Continuous Spatio-Temporal Video Super-Resolution via 2D Gaussian Splatting
por: Shi, Mingyu, et al.
Publicado: (2026)
por: Shi, Mingyu, et al.
Publicado: (2026)
Towards Long-Form Spatio-Temporal Video Grounding
por: Gu, Xin, et al.
Publicado: (2026)
por: Gu, Xin, et al.
Publicado: (2026)
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
por: Aparcedo, Alejandro, et al.
Publicado: (2026)
por: Aparcedo, Alejandro, et al.
Publicado: (2026)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
por: Ahmad, Ghazi Shazan, et al.
Publicado: (2025)
por: Ahmad, Ghazi Shazan, et al.
Publicado: (2025)
Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
por: Choi, Seung hee, et al.
Publicado: (2026)
por: Choi, Seung hee, et al.
Publicado: (2026)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
por: Yu, Li, et al.
Publicado: (2025)
por: Yu, Li, et al.
Publicado: (2025)
SalFoM: Dynamic Saliency Prediction with Video Foundation Models
por: Moradi, Morteza, et al.
Publicado: (2024)
por: Moradi, Morteza, et al.
Publicado: (2024)
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
por: Xiong, Junwen, et al.
Publicado: (2024)
por: Xiong, Junwen, et al.
Publicado: (2024)
Ejemplares similares
-
EditIQ: Automated Cinematic Editing of Static Wide-Angle Videos via Dialogue Interpretation and Saliency Cues
por: Girmaji, Rohit, et al.
Publicado: (2025) -
Simplifying Knowledge Transfer in Pretrained Models
por: Jain, Siddharth, et al.
Publicado: (2025) -
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
por: Moradi, Morteza, et al.
Publicado: (2024) -
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
por: Wang, Yubin, et al.
Publicado: (2024) -
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
por: Siddiqui, Nyle, et al.
Publicado: (2025)