LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression Comprehension
Fuente:
arXiv
Saved in:
| Main Authors: | Cardiel, Amaia, Zablocki, Eloi, Ramzi, Elias, Siméoni, Oriane, Cord, Matthieu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GIFT: A Framework Towards Global Interpretable Faithful Textual Explanations of Vision Classifiers
by: Zablocki, Éloi, et al.
Published: (2024)
by: Zablocki, Éloi, et al.
Published: (2024)
Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models
by: Wysoczańska, Monika, et al.
Published: (2024)
by: Wysoczańska, Monika, et al.
Published: (2024)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
by: Siméoni, Oriane, et al.
Published: (2023)
by: Siméoni, Oriane, et al.
Published: (2023)
DRIV-EX: Counterfactual Explanations for Driving LLMs
by: Cardiel, Amaia, et al.
Published: (2026)
by: Cardiel, Amaia, et al.
Published: (2026)
MAD: Motion Appearance Decoupling for efficient Driving World Models
by: Rahimi, Ahmad, et al.
Published: (2026)
by: Rahimi, Ahmad, et al.
Published: (2026)
GaussRender: Learning 3D Occupancy with Gaussian Rendering
by: Chambon, Loïck, et al.
Published: (2025)
by: Chambon, Loïck, et al.
Published: (2025)
Annealed Winner-Takes-All for Motion Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
ReGentS: Real-World Safety-Critical Driving Scenario Generation Made Stable
by: Yin, Yuan, et al.
Published: (2024)
by: Yin, Yuan, et al.
Published: (2024)
PointBeV: A Sparse Approach to BeV Predictions
by: Chambon, Loick, et al.
Published: (2023)
by: Chambon, Loick, et al.
Published: (2023)
NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
by: Chambon, Loick, et al.
Published: (2025)
by: Chambon, Loick, et al.
Published: (2025)
PPT: Pretraining with Pseudo-Labeled Trajectories for Motion Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
UniTraj: A Unified Framework for Scalable Vehicle Trajectory Prediction
by: Feng, Lan, et al.
Published: (2024)
by: Feng, Lan, et al.
Published: (2024)
MILAN: Milli-Annotations for Lidar Semantic Segmentation
by: Samet, Nermin, et al.
Published: (2024)
by: Samet, Nermin, et al.
Published: (2024)
Valeo4Cast: A Modular Approach to End-to-End Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?
by: Xu, Yihong, et al.
Published: (2023)
by: Xu, Yihong, et al.
Published: (2023)
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
by: Gidaris, Spyros, et al.
Published: (2023)
by: Gidaris, Spyros, et al.
Published: (2023)
RAP: 3D Rasterization Augmented End-to-End Planning
by: Feng, Lan, et al.
Published: (2025)
by: Feng, Lan, et al.
Published: (2025)
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
by: Marouani, Alexis, et al.
Published: (2026)
by: Marouani, Alexis, et al.
Published: (2026)
Federated Black-Box Adaptation for Semantic Segmentation
by: Paranjape, Jay N., et al.
Published: (2024)
by: Paranjape, Jay N., et al.
Published: (2024)
Three Pillars improving Vision Foundation Model Distillation for Lidar
by: Puy, Gilles, et al.
Published: (2023)
by: Puy, Gilles, et al.
Published: (2023)
Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation
by: Vobecky, Antonin, et al.
Published: (2022)
by: Vobecky, Antonin, et al.
Published: (2022)
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
by: Bartoccioni, Florent, et al.
Published: (2025)
by: Bartoccioni, Florent, et al.
Published: (2025)
GalLoP: Learning Global and Local Prompts for Vision-Language Models
by: Lafon, Marc, et al.
Published: (2024)
by: Lafon, Marc, et al.
Published: (2024)
Test-Time Hinting for Black-Box Vision-Language Models
by: Hou, Kaihua, et al.
Published: (2026)
by: Hou, Kaihua, et al.
Published: (2026)
Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation
by: Zhou, Zikun, et al.
Published: (2024)
by: Zhou, Zikun, et al.
Published: (2024)
DistractMIA: Black-Box Membership Inference on Vision-Language Models via Semantic Distraction
by: Tang, Hongyi, et al.
Published: (2026)
by: Tang, Hongyi, et al.
Published: (2026)
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
by: Wysoczańska, Monika, et al.
Published: (2023)
by: Wysoczańska, Monika, et al.
Published: (2023)
Transferable Adversarial Attacks on Black-Box Vision-Language Models
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
Language Models as Black-Box Optimizers for Vision-Language Models
by: Liu, Shihong, et al.
Published: (2023)
by: Liu, Shihong, et al.
Published: (2023)
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
by: Vobecky, Antonin, et al.
Published: (2024)
by: Vobecky, Antonin, et al.
Published: (2024)
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
by: Wang, Lu, et al.
Published: (2025)
by: Wang, Lu, et al.
Published: (2025)
How to Determine the Preferred Image Distribution of a Black-Box Vision-Language Model?
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
Negation-Aware Test-Time Adaptation for Vision-Language Models
by: Han, Haochen, et al.
Published: (2025)
by: Han, Haochen, et al.
Published: (2025)
Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
by: Li, Ling, et al.
Published: (2026)
by: Li, Ling, et al.
Published: (2026)
Referring Expression Comprehension for Small Objects
by: Goto, Kanoko, et al.
Published: (2025)
by: Goto, Kanoko, et al.
Published: (2025)
Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
by: Chen, Jierun, et al.
Published: (2024)
by: Chen, Jierun, et al.
Published: (2024)
Single-Sample Black-Box Membership Inference Attack against Vision-Language Models via Cross-modal Semantic Alignment
by: Li, Jiaqing, et al.
Published: (2026)
by: Li, Jiaqing, et al.
Published: (2026)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Robust Adaptation of Foundation Models with Black-Box Visual Prompting
by: Oh, Changdae, et al.
Published: (2024)
by: Oh, Changdae, et al.
Published: (2024)
Similar Items
-
GIFT: A Framework Towards Global Interpretable Faithful Textual Explanations of Vision Classifiers
by: Zablocki, Éloi, et al.
Published: (2024) -
Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models
by: Wysoczańska, Monika, et al.
Published: (2024) -
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
by: Siméoni, Oriane, et al.
Published: (2023) -
DRIV-EX: Counterfactual Explanations for Driving LLMs
by: Cardiel, Amaia, et al.
Published: (2026) -
MAD: Motion Appearance Decoupling for efficient Driving World Models
by: Rahimi, Ahmad, et al.
Published: (2026)