Training-free Online Video Step Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Zanella, Luca, Mancini, Massimiliano, Wang, Yiming, Tonioni, Alessio, Ricci, Elisa |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Harnessing Large Language Models for Training-free Video Anomaly Detection
by: Zanella, Luca, et al.
Published: (2024)
by: Zanella, Luca, et al.
Published: (2024)
Can Text-to-Video Generation help Video-Language Alignment?
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
by: Das, Deepayan, et al.
Published: (2025)
by: Das, Deepayan, et al.
Published: (2025)
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
by: Das, Deepayan, et al.
Published: (2024)
by: Das, Deepayan, et al.
Published: (2024)
Vocabulary-free Image Classification
by: Conti, Alessandro, et al.
Published: (2023)
by: Conti, Alessandro, et al.
Published: (2023)
Compositional Caching for Training-free Open-vocabulary Attribute Detection
by: Garosi, Marco, et al.
Published: (2025)
by: Garosi, Marco, et al.
Published: (2025)
Vocabulary-free Image Classification and Semantic Segmentation
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
On Large Multimodal Models as Open-World Image Classifiers
by: Conti, Alessandro, et al.
Published: (2025)
by: Conti, Alessandro, et al.
Published: (2025)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
by: Caldarella, Simone, et al.
Published: (2024)
by: Caldarella, Simone, et al.
Published: (2024)
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Large Multimodal Models as General In-Context Classifiers
by: Garosi, Marco, et al.
Published: (2026)
by: Garosi, Marco, et al.
Published: (2026)
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
by: Farina, Matteo, et al.
Published: (2025)
by: Farina, Matteo, et al.
Published: (2025)
Unlearning Personal Data from a Single Image
by: De Min, Thomas, et al.
Published: (2024)
by: De Min, Thomas, et al.
Published: (2024)
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
by: De Min, Thomas, et al.
Published: (2026)
by: De Min, Thomas, et al.
Published: (2026)
Text-Conditioned Resampler For Long Form Video Understanding
by: Korbar, Bruno, et al.
Published: (2023)
by: Korbar, Bruno, et al.
Published: (2023)
MULTIFLOW: Shifting Towards Task-Agnostic Vision-Language Pruning
by: Farina, Matteo, et al.
Published: (2024)
by: Farina, Matteo, et al.
Published: (2024)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
by: Berasi, Davide, et al.
Published: (2025)
by: Berasi, Davide, et al.
Published: (2025)
Frustratingly Easy Test-Time Adaptation of Vision-Language Models
by: Farina, Matteo, et al.
Published: (2024)
by: Farina, Matteo, et al.
Published: (2024)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
by: Plizzari, Chiara, et al.
Published: (2025)
by: Plizzari, Chiara, et al.
Published: (2025)
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
Test-time Vocabulary Adaptation for Language-driven Object Detection
by: Liu, Mingxuan, et al.
Published: (2025)
by: Liu, Mingxuan, et al.
Published: (2025)
From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition
by: Gentile, Francesco, et al.
Published: (2026)
by: Gentile, Francesco, et al.
Published: (2026)
Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers
by: Guimard, Quentin, et al.
Published: (2025)
by: Guimard, Quentin, et al.
Published: (2025)
Less is more: Summarizing Patch Tokens for efficient Multi-Label Class-Incremental Learning
by: De Min, Thomas, et al.
Published: (2024)
by: De Min, Thomas, et al.
Published: (2024)
ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
by: Liberatori, Benedetta, et al.
Published: (2025)
by: Liberatori, Benedetta, et al.
Published: (2025)
Vision-by-Language for Training-Free Compositional Image Retrieval
by: Karthik, Shyamgopal, et al.
Published: (2023)
by: Karthik, Shyamgopal, et al.
Published: (2023)
Retrieval-enriched zero-shot image classification in low-resource domains
by: Dall'Asen, Nicola, et al.
Published: (2024)
by: Dall'Asen, Nicola, et al.
Published: (2024)
Zero-Shot Styled Text Image Generation, but Make It Autoregressive
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
by: Simsar, Enis, et al.
Published: (2024)
by: Simsar, Enis, et al.
Published: (2024)
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
by: Simsar, Enis, et al.
Published: (2023)
by: Simsar, Enis, et al.
Published: (2023)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
by: Guimard, Quentin, et al.
Published: (2026)
by: Guimard, Quentin, et al.
Published: (2026)
Training-free Video Temporal Grounding using Large-scale Pre-trained Models
by: Zheng, Minghang, et al.
Published: (2024)
by: Zheng, Minghang, et al.
Published: (2024)
Specificity-aware reinforcement learning for fine-grained open-world classification
by: Angheben, Samuele, et al.
Published: (2026)
by: Angheben, Samuele, et al.
Published: (2026)
Test-Time Zero-Shot Temporal Action Localization
by: Liberatori, Benedetta, et al.
Published: (2024)
by: Liberatori, Benedetta, et al.
Published: (2024)
Novel class discovery meets foundation models for 3D semantic segmentation
by: Riz, Luigi, et al.
Published: (2023)
by: Riz, Luigi, et al.
Published: (2023)
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025)
by: Vogel, Felix, et al.
Published: (2025)
OnlineHMR: Video-based Online World-Grounded Human Mesh Recovery
by: Zhao, Yiwen, et al.
Published: (2026)
by: Zhao, Yiwen, et al.
Published: (2026)
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
by: Bill, Eric Tillmann, et al.
Published: (2026)
by: Bill, Eric Tillmann, et al.
Published: (2026)
Autoregressive Styled Text Image Generation, but Make it Reliable
by: Zaccagnino, Carmine, et al.
Published: (2025)
by: Zaccagnino, Carmine, et al.
Published: (2025)
Similar Items
-
Harnessing Large Language Models for Training-free Video Anomaly Detection
by: Zanella, Luca, et al.
Published: (2024) -
Can Text-to-Video Generation help Video-Language Alignment?
by: Zanella, Luca, et al.
Published: (2025) -
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
by: Das, Deepayan, et al.
Published: (2025) -
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
by: Das, Deepayan, et al.
Published: (2024) -
Vocabulary-free Image Classification
by: Conti, Alessandro, et al.
Published: (2023)