What does CLIP know about peeling a banana?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cuttano, Claudia, Rosi, Gabriele, Trivigno, Gabriele, Averta, Giuseppe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
von: Cuttano, Claudia, et al.
Veröffentlicht: (2024)
von: Cuttano, Claudia, et al.
Veröffentlicht: (2024)
SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation
von: Cuttano, Claudia, et al.
Veröffentlicht: (2025)
von: Cuttano, Claudia, et al.
Veröffentlicht: (2025)
The revenge of BiSeNet: Efficient Multi-Task Image Segmentation
von: Rosi, Gabriele, et al.
Veröffentlicht: (2024)
von: Rosi, Gabriele, et al.
Veröffentlicht: (2024)
MARCO: Navigating the Unseen Space of Semantic Correspondence
von: Cuttano, Claudia, et al.
Veröffentlicht: (2026)
von: Cuttano, Claudia, et al.
Veröffentlicht: (2026)
PEM: Prototype-based Efficient MaskFormer for Image Segmentation
von: Cavagnero, Niccolò, et al.
Veröffentlicht: (2024)
von: Cavagnero, Niccolò, et al.
Veröffentlicht: (2024)
INSID3: Training-Free In-Context Segmentation with DINOv3
von: Cuttano, Claudia, et al.
Veröffentlicht: (2026)
von: Cuttano, Claudia, et al.
Veröffentlicht: (2026)
Cross-Domain Transfer Learning with CoRTe: Consistent and Reliable Transfer from Black-Box to Lightweight Segmentation Model
von: Cuttano, Claudia, et al.
Veröffentlicht: (2024)
von: Cuttano, Claudia, et al.
Veröffentlicht: (2024)
Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
von: Ramos, Ryan, et al.
Veröffentlicht: (2025)
von: Ramos, Ryan, et al.
Veröffentlicht: (2025)
JIST: Joint Image and Sequence Training for Sequential Visual Place Recognition
von: Berton, Gabriele, et al.
Veröffentlicht: (2024)
von: Berton, Gabriele, et al.
Veröffentlicht: (2024)
To Match or Not to Match: Revisiting Image Matching for Reliable Visual Place Recognition
von: Sferrazza, Davide, et al.
Veröffentlicht: (2025)
von: Sferrazza, Davide, et al.
Veröffentlicht: (2025)
AMEGO: Active Memory from long EGOcentric videos
von: Goletto, Gabriele, et al.
Veröffentlicht: (2024)
von: Goletto, Gabriele, et al.
Veröffentlicht: (2024)
Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic Segmentation
von: Rosi, Gabriele, et al.
Veröffentlicht: (2025)
von: Rosi, Gabriele, et al.
Veröffentlicht: (2025)
EarthMatch: Iterative Coregistration for Fine-grained Localization of Astronaut Photography
von: Berton, Gabriele, et al.
Veröffentlicht: (2024)
von: Berton, Gabriele, et al.
Veröffentlicht: (2024)
The Unreasonable Effectiveness of Pre-Trained Features for Camera Pose Refinement
von: Trivigno, Gabriele, et al.
Veröffentlicht: (2024)
von: Trivigno, Gabriele, et al.
Veröffentlicht: (2024)
Collaborative Visual Place Recognition through Federated Learning
von: Dutto, Mattia, et al.
Veröffentlicht: (2024)
von: Dutto, Mattia, et al.
Veröffentlicht: (2024)
Egocentric zone-aware action recognition across environments
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2024)
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2024)
PrAda: Few-Shot Visual Adaptation for Text-Prompted Segmentation
von: Rosi, Gabriele, et al.
Veröffentlicht: (2026)
von: Rosi, Gabriele, et al.
Veröffentlicht: (2026)
HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2025)
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2025)
FORESCENE: FOREcasting human activity via latent SCENE graphs diffusion
von: Alliegro, Antonio, et al.
Veröffentlicht: (2025)
von: Alliegro, Antonio, et al.
Veröffentlicht: (2025)
HiERO-StepG @ Ego4D Step Grounding Challenge: hierarchical activity understanding enables zero-shot step grounding
von: Zenotto, Andrea, et al.
Veröffentlicht: (2026)
von: Zenotto, Andrea, et al.
Veröffentlicht: (2026)
OpenCity3D: What do Vision-Language Models know about Urban Environments?
von: Bieri, Valentin, et al.
Veröffentlicht: (2025)
von: Bieri, Valentin, et al.
Veröffentlicht: (2025)
Hier-EgoPack: Hierarchical Egocentric Video Understanding with Diverse Task Perspectives
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2025)
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2025)
Learning reusable concepts across different egocentric video understanding tasks
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2025)
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2025)
Scale-Free Image Keypoints Using Differentiable Persistent Homology
von: Barbarani, Giovanni, et al.
Veröffentlicht: (2024)
von: Barbarani, Giovanni, et al.
Veröffentlicht: (2024)
A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2024)
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2024)
Biologically-inspired Semi-supervised Semantic Segmentation for Biomedical Imaging
von: Ciampi, Luca, et al.
Veröffentlicht: (2024)
von: Ciampi, Luca, et al.
Veröffentlicht: (2024)
Semi-Supervised Biomedical Image Segmentation via Diffusion Models and Teacher-Student Co-Training
von: Ciampi, Luca, et al.
Veröffentlicht: (2025)
von: Ciampi, Luca, et al.
Veröffentlicht: (2025)
CA3D: Convolutional-Attentional 3D Nets for Efficient Video Activity Recognition on the Edge
von: Lagani, Gabriele, et al.
Veröffentlicht: (2025)
von: Lagani, Gabriele, et al.
Veröffentlicht: (2025)
Domain Generalization using Action Sequences for Egocentric Action Recognition
von: Nasirimajd, Amirshayan, et al.
Veröffentlicht: (2025)
von: Nasirimajd, Amirshayan, et al.
Veröffentlicht: (2025)
Fixed External Cameras as Common Prior Maps for Active 3D Scene Graph Generation
von: Modi, Giorgia, et al.
Veröffentlicht: (2026)
von: Modi, Giorgia, et al.
Veröffentlicht: (2026)
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots
von: Modi, Giorgia, et al.
Veröffentlicht: (2026)
von: Modi, Giorgia, et al.
Veröffentlicht: (2026)
Comparison of Different Deep Neural Network Models in the Cultural Heritage Domain
von: Boyadzhiev, Teodor, et al.
Veröffentlicht: (2025)
von: Boyadzhiev, Teodor, et al.
Veröffentlicht: (2025)
Finetuning CLIP to Reason about Pairwise Differences
von: Sam, Dylan, et al.
Veröffentlicht: (2024)
von: Sam, Dylan, et al.
Veröffentlicht: (2024)
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
von: Pei, Gensheng, et al.
Veröffentlicht: (2025)
von: Pei, Gensheng, et al.
Veröffentlicht: (2025)
MegaLoc: One Retrieval to Place Them All
von: Berton, Gabriele, et al.
Veröffentlicht: (2025)
von: Berton, Gabriele, et al.
Veröffentlicht: (2025)
How (Mis)calibrated is Your Federated CLIP and What To Do About It?
von: Singha, Mainak, et al.
Veröffentlicht: (2025)
von: Singha, Mainak, et al.
Veröffentlicht: (2025)
Your ViT is Secretly an Image Segmentation Model
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
RAVE: Rate-Adaptive Visual Encoding for 3D Gaussian Splatting
von: Tran, Hoang-Nhat, et al.
Veröffentlicht: (2025)
von: Tran, Hoang-Nhat, et al.
Veröffentlicht: (2025)
What does really matter in image goal navigation?
von: Monaci, Gianluca, et al.
Veröffentlicht: (2025)
von: Monaci, Gianluca, et al.
Veröffentlicht: (2025)
Transient Fault Tolerant Semantic Segmentation for Autonomous Driving
von: Iurada, Leonardo, et al.
Veröffentlicht: (2024)
von: Iurada, Leonardo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
von: Cuttano, Claudia, et al.
Veröffentlicht: (2024) -
SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation
von: Cuttano, Claudia, et al.
Veröffentlicht: (2025) -
The revenge of BiSeNet: Efficient Multi-Task Image Segmentation
von: Rosi, Gabriele, et al.
Veröffentlicht: (2024) -
MARCO: Navigating the Unseen Space of Semantic Correspondence
von: Cuttano, Claudia, et al.
Veröffentlicht: (2026) -
PEM: Prototype-based Efficient MaskFormer for Image Segmentation
von: Cavagnero, Niccolò, et al.
Veröffentlicht: (2024)