ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
Fuente:
arXiv
Saved in:
| Main Authors: | Liberatori, Benedetta, Conti, Alessandro, Vaquero, Lorenzo, Wang, Yiming, Ricci, Elisa, Rota, Paolo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-Shot Temporal Action Localization Through Textual Guidance
by: Liberatori, Benedetta, et al.
Published: (2026)
by: Liberatori, Benedetta, et al.
Published: (2026)
Test-Time Zero-Shot Temporal Action Localization
by: Liberatori, Benedetta, et al.
Published: (2024)
by: Liberatori, Benedetta, et al.
Published: (2024)
Vocabulary-free Image Classification and Semantic Segmentation
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Text-Enhanced Zero-Shot Action Recognition: A training-free approach
by: Bosetti, Massimo, et al.
Published: (2024)
by: Bosetti, Massimo, et al.
Published: (2024)
Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection
by: Tonini, Francesco, et al.
Published: (2025)
by: Tonini, Francesco, et al.
Published: (2025)
Vocabulary-free Image Classification
by: Conti, Alessandro, et al.
Published: (2023)
by: Conti, Alessandro, et al.
Published: (2023)
On Large Multimodal Models as Open-World Image Classifiers
by: Conti, Alessandro, et al.
Published: (2025)
by: Conti, Alessandro, et al.
Published: (2025)
Towards Unconstrained Human-Object Interaction
by: Tonini, Francesco, et al.
Published: (2026)
by: Tonini, Francesco, et al.
Published: (2026)
Conditioned Prompt-Optimization for Continual Deepfake Detection
by: Laiti, Francesco, et al.
Published: (2024)
by: Laiti, Francesco, et al.
Published: (2024)
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Training-Free Semantic Multi-Object Tracking with Vision-Language Models
by: Bonat, Laurence, et al.
Published: (2026)
by: Bonat, Laurence, et al.
Published: (2026)
Dense Motion Captioning
by: Xu, Shiyao, et al.
Published: (2025)
by: Xu, Shiyao, et al.
Published: (2025)
Specificity-aware reinforcement learning for fine-grained open-world classification
by: Angheben, Samuele, et al.
Published: (2026)
by: Angheben, Samuele, et al.
Published: (2026)
From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition
by: Gentile, Francesco, et al.
Published: (2026)
by: Gentile, Francesco, et al.
Published: (2026)
AL-GTD: Deep Active Learning for Gaze Target Detection
by: Tonini, Francesco, et al.
Published: (2024)
by: Tonini, Francesco, et al.
Published: (2024)
Training-free Online Video Step Grounding
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
Large Multimodal Models as General In-Context Classifiers
by: Garosi, Marco, et al.
Published: (2026)
by: Garosi, Marco, et al.
Published: (2026)
Compositional Caching for Training-free Open-vocabulary Attribute Detection
by: Garosi, Marco, et al.
Published: (2025)
by: Garosi, Marco, et al.
Published: (2025)
Can Text-to-Video Generation help Video-Language Alignment?
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
Harnessing Large Language Models for Training-free Video Anomaly Detection
by: Zanella, Luca, et al.
Published: (2024)
by: Zanella, Luca, et al.
Published: (2024)
ViConEx-Med: Visual Concept Explainability via Multi-Concept Token Transformer for Medical Image Analysis
by: Patrício, Cristiano, et al.
Published: (2025)
by: Patrício, Cristiano, et al.
Published: (2025)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
by: Wang, Xiaodong, et al.
Published: (2026)
by: Wang, Xiaodong, et al.
Published: (2026)
GenConViT: Deepfake Video Detection Using Generative Convolutional Vision Transformer
by: Deressa, Deressa Wodajo, et al.
Published: (2023)
by: Deressa, Deressa Wodajo, et al.
Published: (2023)
Retrieval-enriched zero-shot image classification in low-resource domains
by: Dall'Asen, Nicola, et al.
Published: (2024)
by: Dall'Asen, Nicola, et al.
Published: (2024)
Superpowering Open-Vocabulary Object Detectors for X-ray Vision
by: Garcia-Fernandez, Pablo, et al.
Published: (2025)
by: Garcia-Fernandez, Pablo, et al.
Published: (2025)
Human Action Recognition in Still Images Using ConViT
by: Hosseyni, Seyed Rohollah, et al.
Published: (2023)
by: Hosseyni, Seyed Rohollah, et al.
Published: (2023)
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
by: De Min, Thomas, et al.
Published: (2026)
by: De Min, Thomas, et al.
Published: (2026)
SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
by: Sheng, Yuan, et al.
Published: (2025)
by: Sheng, Yuan, et al.
Published: (2025)
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
by: Das, Deepayan, et al.
Published: (2024)
by: Das, Deepayan, et al.
Published: (2024)
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
by: Das, Deepayan, et al.
Published: (2025)
by: Das, Deepayan, et al.
Published: (2025)
Novel class discovery meets foundation models for 3D semantic segmentation
by: Riz, Luigi, et al.
Published: (2023)
by: Riz, Luigi, et al.
Published: (2023)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
by: Lu, Yifan, et al.
Published: (2023)
by: Lu, Yifan, et al.
Published: (2023)
ViStoryBench: Comprehensive Benchmark Suite for Story Visualization
by: Zhuang, Cailin, et al.
Published: (2025)
by: Zhuang, Cailin, et al.
Published: (2025)
ViViD: Video Virtual Try-on using Diffusion Models
by: Fang, Zixun, et al.
Published: (2024)
by: Fang, Zixun, et al.
Published: (2024)
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
by: Tur, Anil Osman, et al.
Published: (2024)
by: Tur, Anil Osman, et al.
Published: (2024)
ConViTac: Aligning Visual-Tactile Fusion with Contrastive Representations
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
Estimation of Psychosocial Work Environment Exposures Through Video Object Detection. Proof of Concept Using CCTV Footage
by: Hansen, Claus D., et al.
Published: (2024)
by: Hansen, Claus D., et al.
Published: (2024)
ViLA: Efficient Video-Language Alignment for Video Question Answering
by: Wang, Xijun, et al.
Published: (2023)
by: Wang, Xijun, et al.
Published: (2023)
SHiNe: Semantic Hierarchy Nexus for Open-vocabulary Object Detection
by: Liu, Mingxuan, et al.
Published: (2024)
by: Liu, Mingxuan, et al.
Published: (2024)
Similar Items
-
Zero-Shot Temporal Action Localization Through Textual Guidance
by: Liberatori, Benedetta, et al.
Published: (2026) -
Test-Time Zero-Shot Temporal Action Localization
by: Liberatori, Benedetta, et al.
Published: (2024) -
Vocabulary-free Image Classification and Semantic Segmentation
by: Conti, Alessandro, et al.
Published: (2024) -
Text-Enhanced Zero-Shot Action Recognition: A training-free approach
by: Bosetti, Massimo, et al.
Published: (2024) -
Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection
by: Tonini, Francesco, et al.
Published: (2025)