Zero-Shot Temporal Action Localization Through Textual Guidance
Fuente:
arXiv
Guardado en:
| Autores principales: | Liberatori, Benedetta, Conti, Alessandro, Vaquero, Lorenzo, Rota, Paolo, Wang, Yiming, Ricci, Elisa |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Test-Time Zero-Shot Temporal Action Localization
por: Liberatori, Benedetta, et al.
Publicado: (2024)
por: Liberatori, Benedetta, et al.
Publicado: (2024)
ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
por: Liberatori, Benedetta, et al.
Publicado: (2025)
por: Liberatori, Benedetta, et al.
Publicado: (2025)
Text-Enhanced Zero-Shot Action Recognition: A training-free approach
por: Bosetti, Massimo, et al.
Publicado: (2024)
por: Bosetti, Massimo, et al.
Publicado: (2024)
Vocabulary-free Image Classification and Semantic Segmentation
por: Conti, Alessandro, et al.
Publicado: (2024)
por: Conti, Alessandro, et al.
Publicado: (2024)
Vocabulary-free Image Classification
por: Conti, Alessandro, et al.
Publicado: (2023)
por: Conti, Alessandro, et al.
Publicado: (2023)
On Large Multimodal Models as Open-World Image Classifiers
por: Conti, Alessandro, et al.
Publicado: (2025)
por: Conti, Alessandro, et al.
Publicado: (2025)
Automatic benchmarking of large multimodal models via iterative experiment programming
por: Conti, Alessandro, et al.
Publicado: (2024)
por: Conti, Alessandro, et al.
Publicado: (2024)
Towards Unconstrained Human-Object Interaction
por: Tonini, Francesco, et al.
Publicado: (2026)
por: Tonini, Francesco, et al.
Publicado: (2026)
Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection
por: Tonini, Francesco, et al.
Publicado: (2025)
por: Tonini, Francesco, et al.
Publicado: (2025)
Conditioned Prompt-Optimization for Continual Deepfake Detection
por: Laiti, Francesco, et al.
Publicado: (2024)
por: Laiti, Francesco, et al.
Publicado: (2024)
Dense Motion Captioning
por: Xu, Shiyao, et al.
Publicado: (2025)
por: Xu, Shiyao, et al.
Publicado: (2025)
Specificity-aware reinforcement learning for fine-grained open-world classification
por: Angheben, Samuele, et al.
Publicado: (2026)
por: Angheben, Samuele, et al.
Publicado: (2026)
OZ-TAL: Online Zero-Shot Temporal Action Localization
por: Han, Chaolei, et al.
Publicado: (2026)
por: Han, Chaolei, et al.
Publicado: (2026)
Training-Free Semantic Multi-Object Tracking with Vision-Language Models
por: Bonat, Laurence, et al.
Publicado: (2026)
por: Bonat, Laurence, et al.
Publicado: (2026)
Towards Completeness: A Generalizable Action Proposal Generator for Zero-Shot Temporal Action Localization
por: Du, Jia-Run, et al.
Publicado: (2024)
por: Du, Jia-Run, et al.
Publicado: (2024)
AL-GTD: Deep Active Learning for Gaze Target Detection
por: Tonini, Francesco, et al.
Publicado: (2024)
por: Tonini, Francesco, et al.
Publicado: (2024)
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
por: Gupta, Akshita, et al.
Publicado: (2024)
por: Gupta, Akshita, et al.
Publicado: (2024)
Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models
por: Han, Chaolei, et al.
Publicado: (2025)
por: Han, Chaolei, et al.
Publicado: (2025)
Large Multimodal Models as General In-Context Classifiers
por: Garosi, Marco, et al.
Publicado: (2026)
por: Garosi, Marco, et al.
Publicado: (2026)
Compositional Caching for Training-free Open-vocabulary Attribute Detection
por: Garosi, Marco, et al.
Publicado: (2025)
por: Garosi, Marco, et al.
Publicado: (2025)
Zero-Shot Personalization of Objects via Textual Inversion
por: Roy, Aniket, et al.
Publicado: (2026)
por: Roy, Aniket, et al.
Publicado: (2026)
iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
por: Agnolucci, Lorenzo, et al.
Publicado: (2024)
por: Agnolucci, Lorenzo, et al.
Publicado: (2024)
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
por: Tur, Anil Osman, et al.
Publicado: (2024)
por: Tur, Anil Osman, et al.
Publicado: (2024)
From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition
por: Gentile, Francesco, et al.
Publicado: (2026)
por: Gentile, Francesco, et al.
Publicado: (2026)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
por: Huang, Wei-Jhe, et al.
Publicado: (2024)
por: Huang, Wei-Jhe, et al.
Publicado: (2024)
Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
por: Song, Yeji, et al.
Publicado: (2024)
por: Song, Yeji, et al.
Publicado: (2024)
Zero-Shot Textual Explanations via Translating Decision-Critical Features
por: Yamauchi, Toshinori, et al.
Publicado: (2025)
por: Yamauchi, Toshinori, et al.
Publicado: (2025)
Zero-Shot Temporal Interaction Localization for Egocentric Videos
por: Zhang, Erhang, et al.
Publicado: (2025)
por: Zhang, Erhang, et al.
Publicado: (2025)
Zero-Shot Faithful Textual Explanations via Directional-Derivative Influence on Predictions
por: Yamauchi, Toshinori, et al.
Publicado: (2026)
por: Yamauchi, Toshinori, et al.
Publicado: (2026)
Zero-Shot Visual Concept Blending Without Text Guidance
por: Makino, Hiroya, et al.
Publicado: (2025)
por: Makino, Hiroya, et al.
Publicado: (2025)
Text-guided Zero-Shot Object Localization
por: Wang, Jingjing, et al.
Publicado: (2024)
por: Wang, Jingjing, et al.
Publicado: (2024)
Retrieval-enriched zero-shot image classification in low-resource domains
por: Dall'Asen, Nicola, et al.
Publicado: (2024)
por: Dall'Asen, Nicola, et al.
Publicado: (2024)
Novel Semantic Prompting for Zero-Shot Action Recognition
por: Iqbal, Salman, et al.
Publicado: (2026)
por: Iqbal, Salman, et al.
Publicado: (2026)
Continual Learning Improves Zero-Shot Action Recognition
por: Gowda, Shreyank N, et al.
Publicado: (2024)
por: Gowda, Shreyank N, et al.
Publicado: (2024)
Zero-Shot Interpretable Image Steganalysis for Invertible Image Hiding
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
por: Lin, Haoqiang, et al.
Publicado: (2025)
por: Lin, Haoqiang, et al.
Publicado: (2025)
SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization
por: Borse, Shubhankar, et al.
Publicado: (2025)
por: Borse, Shubhankar, et al.
Publicado: (2025)
Zero-Shot Action Recognition in Surveillance Videos
por: Pereira, Joao, et al.
Publicado: (2024)
por: Pereira, Joao, et al.
Publicado: (2024)
Superpowering Open-Vocabulary Object Detectors for X-ray Vision
por: Garcia-Fernandez, Pablo, et al.
Publicado: (2025)
por: Garcia-Fernandez, Pablo, et al.
Publicado: (2025)
Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation
por: Liu, Ting, et al.
Publicado: (2025)
por: Liu, Ting, et al.
Publicado: (2025)
Ejemplares similares
-
Test-Time Zero-Shot Temporal Action Localization
por: Liberatori, Benedetta, et al.
Publicado: (2024) -
ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
por: Liberatori, Benedetta, et al.
Publicado: (2025) -
Text-Enhanced Zero-Shot Action Recognition: A training-free approach
por: Bosetti, Massimo, et al.
Publicado: (2024) -
Vocabulary-free Image Classification and Semantic Segmentation
por: Conti, Alessandro, et al.
Publicado: (2024) -
Vocabulary-free Image Classification
por: Conti, Alessandro, et al.
Publicado: (2023)