Saved in:
| Main Authors: | Angheben, Samuele, Berasi, Davide, Conti, Alessandro, Ricci, Elisa, Wang, Yiming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.03197 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
by: Berasi, Davide, et al.
Published: (2025)
by: Berasi, Davide, et al.
Published: (2025)
Test-Time Zero-Shot Temporal Action Localization
by: Liberatori, Benedetta, et al.
Published: (2024)
by: Liberatori, Benedetta, et al.
Published: (2024)
Zero-Shot Temporal Action Localization Through Textual Guidance
by: Liberatori, Benedetta, et al.
Published: (2026)
by: Liberatori, Benedetta, et al.
Published: (2026)
Vocabulary-free Image Classification and Semantic Segmentation
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Vocabulary-free Image Classification
by: Conti, Alessandro, et al.
Published: (2023)
by: Conti, Alessandro, et al.
Published: (2023)
ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
by: Liberatori, Benedetta, et al.
Published: (2025)
by: Liberatori, Benedetta, et al.
Published: (2025)
On Large Multimodal Models as Open-World Image Classifiers
by: Conti, Alessandro, et al.
Published: (2025)
by: Conti, Alessandro, et al.
Published: (2025)
Is CLIP the main roadblock for fine-grained open-world perception?
by: Bianchi, Lorenzo, et al.
Published: (2024)
by: Bianchi, Lorenzo, et al.
Published: (2024)
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
by: Tur, Anil Osman, et al.
Published: (2024)
by: Tur, Anil Osman, et al.
Published: (2024)
Retrieval-enriched zero-shot image classification in low-resource domains
by: Dall'Asen, Nicola, et al.
Published: (2024)
by: Dall'Asen, Nicola, et al.
Published: (2024)
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Large Multimodal Models as General In-Context Classifiers
by: Garosi, Marco, et al.
Published: (2026)
by: Garosi, Marco, et al.
Published: (2026)
Towards Unconstrained Human-Object Interaction
by: Tonini, Francesco, et al.
Published: (2026)
by: Tonini, Francesco, et al.
Published: (2026)
Compositional Caching for Training-free Open-vocabulary Attribute Detection
by: Garosi, Marco, et al.
Published: (2025)
by: Garosi, Marco, et al.
Published: (2025)
Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection
by: Tonini, Francesco, et al.
Published: (2025)
by: Tonini, Francesco, et al.
Published: (2025)
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
by: Das, Deepayan, et al.
Published: (2024)
by: Das, Deepayan, et al.
Published: (2024)
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
by: Das, Deepayan, et al.
Published: (2025)
by: Das, Deepayan, et al.
Published: (2025)
Local Foreground Selection aware Attentive Feature Reconstruction for few-shot fine-grained plant species classification
by: Zulfiqar, Aisha, et al.
Published: (2025)
by: Zulfiqar, Aisha, et al.
Published: (2025)
Harnessing Large Language Models for Training-free Video Anomaly Detection
by: Zanella, Luca, et al.
Published: (2024)
by: Zanella, Luca, et al.
Published: (2024)
Novel class discovery meets foundation models for 3D semantic segmentation
by: Riz, Luigi, et al.
Published: (2023)
by: Riz, Luigi, et al.
Published: (2023)
Training-free Online Video Step Grounding
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
How to Take a Memorable Picture? Empowering Users with Actionable Feedback
by: Laiti, Francesco, et al.
Published: (2026)
by: Laiti, Francesco, et al.
Published: (2026)
The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understanding
by: Bianchi, Lorenzo, et al.
Published: (2023)
by: Bianchi, Lorenzo, et al.
Published: (2023)
Bio-inspired fine-tuning for selective transfer learning in image classification
by: Davila, Ana, et al.
Published: (2026)
by: Davila, Ana, et al.
Published: (2026)
Democratizing Fine-grained Visual Recognition with Large Language Models
by: Liu, Mingxuan, et al.
Published: (2024)
by: Liu, Mingxuan, et al.
Published: (2024)
Can Text-to-Video Generation help Video-Language Alignment?
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
Performance of computer vision algorithms for fine-grained classification using crowdsourced insect images
by: Pucci, Rita, et al.
Published: (2024)
by: Pucci, Rita, et al.
Published: (2024)
Location embedding based pairwise distance learning for fine-grained diagnosis of urinary stones
by: Jin, Qiangguo, et al.
Published: (2024)
by: Jin, Qiangguo, et al.
Published: (2024)
Comparison of fine-tuning strategies for transfer learning in medical image classification
by: Davila, Ana, et al.
Published: (2024)
by: Davila, Ana, et al.
Published: (2024)
Demographic-aware fine-grained visual recognition of pediatric wrist pathologies
by: Ahmed, Ammar, et al.
Published: (2025)
by: Ahmed, Ammar, et al.
Published: (2025)
Incremental Object-Based Novelty Detection with Feedback Loop
by: Caldarella, Simone, et al.
Published: (2023)
by: Caldarella, Simone, et al.
Published: (2023)
Adaptive receptive field-based spatial-frequency feature reconstruction network for few-shot fine-grained image classification
by: Zhang, Linyue, et al.
Published: (2026)
by: Zhang, Linyue, et al.
Published: (2026)
Hairmony: Fairness-aware hairstyle classification
by: Meishvili, Givi, et al.
Published: (2024)
by: Meishvili, Givi, et al.
Published: (2024)
An uncertainty-aware Bayesian framework for machine learning classification models: A case study in land cover classification
by: Bilson, Samuel, et al.
Published: (2025)
by: Bilson, Samuel, et al.
Published: (2025)
Infusing fine-grained visual knowledge to Vision-Language Models
by: Ypsilantis, Nikolaos-Antonios, et al.
Published: (2025)
by: Ypsilantis, Nikolaos-Antonios, et al.
Published: (2025)
Plant identification in an open-world (LifeCLEF 2016)
by: Goeau, Herve, et al.
Published: (2025)
by: Goeau, Herve, et al.
Published: (2025)
Training-Free Semantic Multi-Object Tracking with Vision-Language Models
by: Bonat, Laurence, et al.
Published: (2026)
by: Bonat, Laurence, et al.
Published: (2026)
3D Object Detection from Images for Autonomous Driving: A Survey
by: Ma, Xinzhu, et al.
Published: (2022)
by: Ma, Xinzhu, et al.
Published: (2022)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
by: Caldarella, Simone, et al.
Published: (2024)
by: Caldarella, Simone, et al.
Published: (2024)
Supervised contrastive learning for cell stage classification of animal embryos
by: Hachani, Yasmine, et al.
Published: (2025)
by: Hachani, Yasmine, et al.
Published: (2025)
Similar Items
-
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
by: Berasi, Davide, et al.
Published: (2025) -
Test-Time Zero-Shot Temporal Action Localization
by: Liberatori, Benedetta, et al.
Published: (2024) -
Zero-Shot Temporal Action Localization Through Textual Guidance
by: Liberatori, Benedetta, et al.
Published: (2026) -
Vocabulary-free Image Classification and Semantic Segmentation
by: Conti, Alessandro, et al.
Published: (2024) -
Vocabulary-free Image Classification
by: Conti, Alessandro, et al.
Published: (2023)