TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911185868161024 |
|---|---|
| author | Kontostathis, Ioannis Apostolidis, Evlampios Mezaris, Vasileios |
| author_facet | Kontostathis, Ioannis Apostolidis, Evlampios Mezaris, Vasileios |
| contents | In this paper, we deal with the task of text-driven saliency detection in 360-degrees videos. For this, we introduce the TSV360 dataset which includes 16,000 triplets of ERP frames, textual descriptions of salient objects/events in these frames, and the associated ground-truth saliency maps. Following, we extend and adapt a SOTA visual-based approach for 360-degrees video saliency detection, and develop the TSalV360 method that takes into account a user-provided text description of the desired objects and/or events. This method leverages a SOTA vision-language model for data representation and integrates a similarity estimation module and a viewport spatio-temporal cross-attention mechanism, to discover dependencies between the different data modalities. Quantitative and qualitative evaluations using the TSV360 dataset, showed the competitiveness of TSalV360 compared to a SOTA visual-based approach and documented its competency to perform customized text-driven saliency detection in 360-degrees videos. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_26208 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos Kontostathis, Ioannis Apostolidis, Evlampios Mezaris, Vasileios Computer Vision and Pattern Recognition In this paper, we deal with the task of text-driven saliency detection in 360-degrees videos. For this, we introduce the TSV360 dataset which includes 16,000 triplets of ERP frames, textual descriptions of salient objects/events in these frames, and the associated ground-truth saliency maps. Following, we extend and adapt a SOTA visual-based approach for 360-degrees video saliency detection, and develop the TSalV360 method that takes into account a user-provided text description of the desired objects and/or events. This method leverages a SOTA vision-language model for data representation and integrates a similarity estimation module and a viewport spatio-temporal cross-attention mechanism, to discover dependencies between the different data modalities. Quantitative and qualitative evaluations using the TSV360 dataset, showed the competitiveness of TSalV360 compared to a SOTA visual-based approach and documented its competency to perform customized text-driven saliency detection in 360-degrees videos. |
| title | TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.26208 |