The Power of One: A Single Example is All it Takes for Segmentation in VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Hossain, Mir Rayat Imtiaz, Siam, Mennatullah, Sigal, Leonid, Little, James J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
Framework-agnostic Semantically-aware Global Reasoning for Segmentation
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2022)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2022)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023)
by: Goyal, Raghav, et al.
Published: (2023)
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
by: Luo, Jiayun, et al.
Published: (2024)
by: Luo, Jiayun, et al.
Published: (2024)
PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?
by: Siam, Mennatullah
Published: (2025)
by: Siam, Mennatullah
Published: (2025)
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
by: Siam, Mennatullah
Published: (2025)
by: Siam, Mennatullah
Published: (2025)
Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving
by: Cheshmi, Leila, et al.
Published: (2025)
by: Cheshmi, Leila, et al.
Published: (2025)
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
by: Salamatian, Ali, et al.
Published: (2025)
by: Salamatian, Ali, et al.
Published: (2025)
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
by: Chou, Shih-Han, et al.
Published: (2024)
by: Chou, Shih-Han, et al.
Published: (2024)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
by: Chou, Shih-Han, et al.
Published: (2023)
by: Chou, Shih-Han, et al.
Published: (2023)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
by: Rahman, Tanzila, et al.
Published: (2026)
by: Rahman, Tanzila, et al.
Published: (2026)
Test-Time Consistency in Vision Language Models
by: Chou, Shih-Han, et al.
Published: (2025)
by: Chou, Shih-Han, et al.
Published: (2025)
Generalized Few-Shot Semantic Segmentation in Remote Sensing: Challenge and Benchmark
by: Broni-Bediako, Clifford, et al.
Published: (2024)
by: Broni-Bediako, Clifford, et al.
Published: (2024)
MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
by: Karim, Rezaul, et al.
Published: (2023)
by: Karim, Rezaul, et al.
Published: (2023)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
by: Chinchure, Aditya, et al.
Published: (2025)
by: Chinchure, Aditya, et al.
Published: (2025)
Dynamics Based Neural Encoding with Inter-Intra Region Connectivity
by: Gamal, Mai, et al.
Published: (2024)
by: Gamal, Mai, et al.
Published: (2024)
Teaching VLMs to Localize Specific Objects from In-context Examples
by: Doveh, Sivan, et al.
Published: (2024)
by: Doveh, Sivan, et al.
Published: (2024)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
by: Bhatt, Gaurav, et al.
Published: (2024)
by: Bhatt, Gaurav, et al.
Published: (2024)
Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models
by: Luo, Jiayun, et al.
Published: (2023)
by: Luo, Jiayun, et al.
Published: (2023)
Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal Networks
by: Kowal, Matthew, et al.
Published: (2022)
by: Kowal, Matthew, et al.
Published: (2022)
Selecting Fine-Tuning Examples by Quizzing VLMs
by: Ji, Tenghao, et al.
Published: (2025)
by: Ji, Tenghao, et al.
Published: (2025)
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
by: Chandhok, Shivam, et al.
Published: (2024)
by: Chandhok, Shivam, et al.
Published: (2024)
Segment Using Just One Example
by: Vora, Pratik, et al.
Published: (2024)
by: Vora, Pratik, et al.
Published: (2024)
Factorized Video Autoencoders for Efficient Generative Modelling
by: Suhail, Mohammed, et al.
Published: (2024)
by: Suhail, Mohammed, et al.
Published: (2024)
OneVOS: Unifying Video Object Segmentation with All-in-One Transformer Framework
by: Li, Wanyun, et al.
Published: (2024)
by: Li, Wanyun, et al.
Published: (2024)
UnSeg: One Universal Unlearnable Example Generator is Enough against All Image Segmentation
by: Sun, Ye, et al.
Published: (2024)
by: Sun, Ye, et al.
Published: (2024)
Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models
by: Shukla, Pushkar, et al.
Published: (2025)
by: Shukla, Pushkar, et al.
Published: (2025)
A Vision Centric Remote Sensing Benchmark
by: Adejumo, Abduljaleel, et al.
Published: (2025)
by: Adejumo, Abduljaleel, et al.
Published: (2025)
OMG-Seg: Is One Model Good Enough For All Segmentation?
by: Li, Xiangtai, et al.
Published: (2024)
by: Li, Xiangtai, et al.
Published: (2024)
One-shot Optimized Steering Vector for Hallucination Mitigation for VLMs
by: Shi, Youxu, et al.
Published: (2026)
by: Shi, Youxu, et al.
Published: (2026)
The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
by: Azad, Asif, et al.
Published: (2025)
by: Azad, Asif, et al.
Published: (2025)
One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image
by: Shereen, Ezzeldin, et al.
Published: (2025)
by: Shereen, Ezzeldin, et al.
Published: (2025)
Cross-Domain Semantic Segmentation on Inconsistent Taxonomy using VLMs
by: Lim, Jeongkee, et al.
Published: (2024)
by: Lim, Jeongkee, et al.
Published: (2024)
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
by: Sakai, Shunsuke, et al.
Published: (2025)
by: Sakai, Shunsuke, et al.
Published: (2025)
Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
by: Xu, Bicheng, et al.
Published: (2024)
by: Xu, Bicheng, et al.
Published: (2024)
Prompt2Perturb (P2P): Text-Guided Diffusion-Based Adversarial Attacks on Breast Ultrasound Images
by: Medghalchi, Yasamin, et al.
Published: (2024)
by: Medghalchi, Yasamin, et al.
Published: (2024)
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
by: Mahdizadeh, Ailar, et al.
Published: (2026)
by: Mahdizadeh, Ailar, et al.
Published: (2026)
MCP-MedSAM: A Powerful Lightweight Medical Segment Anything Model Trained with a Single GPU in Just One Day
by: Lyu, Donghang, et al.
Published: (2024)
by: Lyu, Donghang, et al.
Published: (2024)
Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
Exploring Prompt Alignment with Clinical Factors in Zero-Shot Segmentation VLMs for NSCLC Tumor Segmentation
by: Pai, Suraj, et al.
Published: (2026)
by: Pai, Suraj, et al.
Published: (2026)
Similar Items
-
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024) -
Framework-agnostic Semantically-aware Global Reasoning for Segmentation
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2022) -
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023) -
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
by: Luo, Jiayun, et al.
Published: (2024) -
PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?
by: Siam, Mennatullah
Published: (2025)