Show or Tell? Effectively prompting Vision-Language Models for semantic segmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Avogaro, Niccolo, Frick, Thomas, Rigotti, Mattia, Bartezzaghi, Andrea, Janicki, Filip, Malossi, Cristiano, Schindler, Konrad, Assaf, Roy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation
por: Avogaro, Niccolo, et al.
Publicado: (2025)
por: Avogaro, Niccolo, et al.
Publicado: (2025)
Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models
por: Farronato, Nicola, et al.
Publicado: (2026)
por: Farronato, Nicola, et al.
Publicado: (2026)
Outline-Guided Object Inpainting with Diffusion Models
por: Pobitzer, Markus, et al.
Publicado: (2024)
por: Pobitzer, Markus, et al.
Publicado: (2024)
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
por: Ebouky, Brown, et al.
Publicado: (2026)
por: Ebouky, Brown, et al.
Publicado: (2026)
Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
por: Ebouky, Brown, et al.
Publicado: (2025)
por: Ebouky, Brown, et al.
Publicado: (2025)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
por: Avogaro, Niccolo, et al.
Publicado: (2026)
por: Avogaro, Niccolo, et al.
Publicado: (2026)
Faster by Design: Interactive Aerodynamics via Neural Surrogates Trained on Expert-Validated CFD
por: Thumiger, Nicholas, et al.
Publicado: (2026)
por: Thumiger, Nicholas, et al.
Publicado: (2026)
Eliciting Reasoning in Language Models with Cognitive Tools
por: Ebouky, Brown, et al.
Publicado: (2025)
por: Ebouky, Brown, et al.
Publicado: (2025)
Q-SAM2: Accurate Quantization for Segment Anything Model 2
por: Farronato, Nicola, et al.
Publicado: (2025)
por: Farronato, Nicola, et al.
Publicado: (2025)
DepthSeg: Depth prompting in remote sensing semantic segmentation
por: Zhou, Ning, et al.
Publicado: (2025)
por: Zhou, Ning, et al.
Publicado: (2025)
Combining Data Generation and Active Learning for Low-Resource Question Answering
por: Kimmich, Maximilian, et al.
Publicado: (2022)
por: Kimmich, Maximilian, et al.
Publicado: (2022)
GIST: Gauge-Invariant Spectral Transformers for Scalable Graph Neural Operators
por: Rigotti, Mattia, et al.
Publicado: (2026)
por: Rigotti, Mattia, et al.
Publicado: (2026)
Novel class discovery meets foundation models for 3D semantic segmentation
por: Riz, Luigi, et al.
Publicado: (2023)
por: Riz, Luigi, et al.
Publicado: (2023)
Diffusion Is Your Friend in Show, Suggest and Tell
por: Hu, Jia Cheng, et al.
Publicado: (2025)
por: Hu, Jia Cheng, et al.
Publicado: (2025)
SMOL-MapSeg: Show Me One Label as prompt
por: Yuan, Yunshuang, et al.
Publicado: (2025)
por: Yuan, Yunshuang, et al.
Publicado: (2025)
Hybridnet for depth estimation and semantic segmentation
por: Sánchez-Escobedo, Dalila, et al.
Publicado: (2024)
por: Sánchez-Escobedo, Dalila, et al.
Publicado: (2024)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
por: Chen, Harold Haodong, et al.
Publicado: (2026)
por: Chen, Harold Haodong, et al.
Publicado: (2026)
Tumor segmentation on whole slide images: training or prompting?
por: Wu, Huaqian, et al.
Publicado: (2024)
por: Wu, Huaqian, et al.
Publicado: (2024)
Shift and matching queries for video semantic segmentation
por: Mizuno, Tsubasa, et al.
Publicado: (2024)
por: Mizuno, Tsubasa, et al.
Publicado: (2024)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
por: Kalluri, Tarun, et al.
Publicado: (2024)
por: Kalluri, Tarun, et al.
Publicado: (2024)
SITUATE: Indoor Human Trajectory Prediction through Geometric Features and Self-Supervised Vision Representation
por: Capogrosso, Luigi, et al.
Publicado: (2024)
por: Capogrosso, Luigi, et al.
Publicado: (2024)
Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic Segmentation
por: Rosi, Gabriele, et al.
Publicado: (2025)
por: Rosi, Gabriele, et al.
Publicado: (2025)
MDiFF: Exploiting Multimodal Score-based Diffusion Models for New Fashion Product Performance Forecasting
por: Avogaro, Andrea, et al.
Publicado: (2024)
por: Avogaro, Andrea, et al.
Publicado: (2024)
Dif4FF: Leveraging Multimodal Diffusion Models and Graph Neural Networks for Accurate New Fashion Product Performance Forecasting
por: Avogaro, Andrea, et al.
Publicado: (2024)
por: Avogaro, Andrea, et al.
Publicado: (2024)
Show, Don't Tell: Detecting Novel Objects by Watching Human Videos
por: Akl, James, et al.
Publicado: (2026)
por: Akl, James, et al.
Publicado: (2026)
SegmentAnyTree: A sensor and platform agnostic deep learning model for tree segmentation using laser scanning data
por: Wielgosz, Maciej, et al.
Publicado: (2024)
por: Wielgosz, Maciej, et al.
Publicado: (2024)
Next day fire prediction via semantic segmentation
por: Alexis, Konstantinos, et al.
Publicado: (2024)
por: Alexis, Konstantinos, et al.
Publicado: (2024)
New Fashion Products Performance Forecasting: A Survey on Evolutions, Models and Emerging Trends
por: Avogaro, Andrea, et al.
Publicado: (2025)
por: Avogaro, Andrea, et al.
Publicado: (2025)
Tree semantic segmentation from aerial image time series
por: Ramesh, Venkatesh, et al.
Publicado: (2024)
por: Ramesh, Venkatesh, et al.
Publicado: (2024)
Impact of LiDAR visualisations on semantic segmentation of archaeological objects
por: Jaturapitpornchai, Raveerat, et al.
Publicado: (2024)
por: Jaturapitpornchai, Raveerat, et al.
Publicado: (2024)
Out-of-distribution data supervision towards biomedical semantic segmentation
por: Gao, Yiquan, et al.
Publicado: (2025)
por: Gao, Yiquan, et al.
Publicado: (2025)
Submodular video object proposal selection for semantic object segmentation
por: Wang, Tinghuai
Publicado: (2024)
por: Wang, Tinghuai
Publicado: (2024)
A Retrospect to Multi-prompt Learning across Vision and Language
por: Chen, Ziliang, et al.
Publicado: (2025)
por: Chen, Ziliang, et al.
Publicado: (2025)
Pseudolabel guided pixels contrast for domain adaptive semantic segmentation
por: Xiang, Jianzi, et al.
Publicado: (2025)
por: Xiang, Jianzi, et al.
Publicado: (2025)
Context-self contrastive pretraining for crop type semantic segmentation
por: Tarasiou, Michail, et al.
Publicado: (2021)
por: Tarasiou, Michail, et al.
Publicado: (2021)
Soft labelling for semantic segmentation: Bringing coherence to label down-sampling
por: Alcover-Couso, Roberto, et al.
Publicado: (2023)
por: Alcover-Couso, Roberto, et al.
Publicado: (2023)
Impact of color and mixing proportion of synthetic point clouds on semantic segmentation
por: Zhou, Shaojie, et al.
Publicado: (2024)
por: Zhou, Shaojie, et al.
Publicado: (2024)
SSR: SAM is a Strong Regularizer for domain adaptive semantic segmentation
por: Ge, Yanqi, et al.
Publicado: (2024)
por: Ge, Yanqi, et al.
Publicado: (2024)
Unsupervised semantic segmentation of urban high-density multispectral point clouds
por: Oinonen, Oona, et al.
Publicado: (2024)
por: Oinonen, Oona, et al.
Publicado: (2024)
Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models
por: Fu, Shuai, et al.
Publicado: (2024)
por: Fu, Shuai, et al.
Publicado: (2024)
Ejemplares similares
-
VP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation
por: Avogaro, Niccolo, et al.
Publicado: (2025) -
Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models
por: Farronato, Nicola, et al.
Publicado: (2026) -
Outline-Guided Object Inpainting with Diffusion Models
por: Pobitzer, Markus, et al.
Publicado: (2024) -
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
por: Ebouky, Brown, et al.
Publicado: (2026) -
Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
por: Ebouky, Brown, et al.
Publicado: (2025)