Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ding, Guodong, Yao, Angela |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Conceptrol: Concept Control of Zero-shot Personalized Image Generation
von: He, Qiyuan, et al.
Veröffentlicht: (2025)
von: He, Qiyuan, et al.
Veröffentlicht: (2025)
Zero-shot Vision-Language Reranking for Cross-View Geolocalization
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026)
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026)
Zero-shot detection of buildings in mobile LiDAR using Language Vision Model
von: Goo, June Moh, et al.
Veröffentlicht: (2024)
von: Goo, June Moh, et al.
Veröffentlicht: (2024)
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
von: Guo, Grace, et al.
Veröffentlicht: (2024)
von: Guo, Grace, et al.
Veröffentlicht: (2024)
Zeus: Zero-shot LLM Instruction for Union Segmentation in Multimodal Medical Imaging
von: Dai, Siyuan, et al.
Veröffentlicht: (2025)
von: Dai, Siyuan, et al.
Veröffentlicht: (2025)
A Vision-Language Foundation Model for Zero-shot Clinical Collaboration and Automated Concept Discovery in Dermatology
von: Yan, Siyuan, et al.
Veröffentlicht: (2026)
von: Yan, Siyuan, et al.
Veröffentlicht: (2026)
P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation
von: Sheng, Kai, et al.
Veröffentlicht: (2026)
von: Sheng, Kai, et al.
Veröffentlicht: (2026)
Hierarchically Robust Zero-shot Vision-language Models
von: Dong, Junhao, et al.
Veröffentlicht: (2026)
von: Dong, Junhao, et al.
Veröffentlicht: (2026)
Enhancing Zero-shot Personalized Image Aesthetics Assessment with Profile-aware Multimodal LLM
von: Wang, Chun, et al.
Veröffentlicht: (2026)
von: Wang, Chun, et al.
Veröffentlicht: (2026)
BootPIG: Bootstrapping Zero-shot Personalized Image Generation Capabilities in Pretrained Diffusion Models
von: Purushwalkam, Senthil, et al.
Veröffentlicht: (2024)
von: Purushwalkam, Senthil, et al.
Veröffentlicht: (2024)
Zero-shot World Models Are Developmentally Efficient Learners
von: Aw, Khai Loong, et al.
Veröffentlicht: (2026)
von: Aw, Khai Loong, et al.
Veröffentlicht: (2026)
Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
InstantID: Zero-shot Identity-Preserving Generation in Seconds
von: Wang, Qixun, et al.
Veröffentlicht: (2024)
von: Wang, Qixun, et al.
Veröffentlicht: (2024)
Ego: Embedding-Guided Personalization of Vision-Language Models
von: Seifi, Soroush, et al.
Veröffentlicht: (2026)
von: Seifi, Soroush, et al.
Veröffentlicht: (2026)
GEASS: Gated Evidence-Adaptive Selective Caption Trust for Vision-Language Models
von: Li, Zeshang, et al.
Veröffentlicht: (2026)
von: Li, Zeshang, et al.
Veröffentlicht: (2026)
AIR: Zero-shot Generative Model Adaptation with Iterative Refinement
von: Liu, Guimeng, et al.
Veröffentlicht: (2025)
von: Liu, Guimeng, et al.
Veröffentlicht: (2025)
OAD-Promoter: Enhancing Zero-shot VQA using Large Language Models with Object Attribute Description
von: Xu, Quanxing, et al.
Veröffentlicht: (2025)
von: Xu, Quanxing, et al.
Veröffentlicht: (2025)
Prompting Large Vision-Language Models for Compositional Reasoning
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
Zoom-shot: Fast and Efficient Unsupervised Zero-Shot Transfer of CLIP to Vision Encoders with Multimodal Loss
von: Shipard, Jordan, et al.
Veröffentlicht: (2024)
von: Shipard, Jordan, et al.
Veröffentlicht: (2024)
Prompt-based Visual Alignment for Zero-shot Policy Transfer
von: Gao, Haihan, et al.
Veröffentlicht: (2024)
von: Gao, Haihan, et al.
Veröffentlicht: (2024)
Continual Learning in Vision-Language Models via Aligned Model Merging
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
Segment-Anything Models Achieve Zero-shot Robustness in Autonomous Driving
von: Yan, Jun, et al.
Veröffentlicht: (2024)
von: Yan, Jun, et al.
Veröffentlicht: (2024)
Visual Space Optimization for Zero-shot Learning
von: Wang, Xinsheng, et al.
Veröffentlicht: (2019)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2019)
Fine-gained Zero-shot Video Sampling
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models
von: Bharadwaj, Sagar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Sagar, et al.
Veröffentlicht: (2026)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2024)
von: An, Ruichuan, et al.
Veröffentlicht: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
In-Context Learning Improves Compositional Understanding of Vision-Language Models
von: Nulli, Matteo, et al.
Veröffentlicht: (2024)
von: Nulli, Matteo, et al.
Veröffentlicht: (2024)
Zero-shot Concept Bottleneck Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers
von: Stanić, Aleksandar, et al.
Veröffentlicht: (2024)
von: Stanić, Aleksandar, et al.
Veröffentlicht: (2024)
YoChameleon: Personalized Vision and Language Generation
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
GAZE:Governance-Aware pre-annotation for Zero-shot World Model Environments
von: Krishna, Leela, et al.
Veröffentlicht: (2025)
von: Krishna, Leela, et al.
Veröffentlicht: (2025)
Jailbreaks on Vision Language Model via Multimodal Reasoning
von: Noheria, Aarush, et al.
Veröffentlicht: (2026)
von: Noheria, Aarush, et al.
Veröffentlicht: (2026)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
von: Zhang, Pu, et al.
Veröffentlicht: (2025)
von: Zhang, Pu, et al.
Veröffentlicht: (2025)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
von: Xu, Zhenlin, et al.
Veröffentlicht: (2023)
von: Xu, Zhenlin, et al.
Veröffentlicht: (2023)
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
von: Bendou, Yassir, et al.
Veröffentlicht: (2024)
von: Bendou, Yassir, et al.
Veröffentlicht: (2024)
AA-CLIP: Enhancing Zero-shot Anomaly Detection via Anomaly-Aware CLIP
von: Ma, Wenxin, et al.
Veröffentlicht: (2025)
von: Ma, Wenxin, et al.
Veröffentlicht: (2025)
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment
von: Kim, Younghyun, et al.
Veröffentlicht: (2025)
von: Kim, Younghyun, et al.
Veröffentlicht: (2025)
LIBERO-X: Robustness Litmus for Vision-Language-Action Models
von: Wang, Guodong, et al.
Veröffentlicht: (2026)
von: Wang, Guodong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Conceptrol: Concept Control of Zero-shot Personalized Image Generation
von: He, Qiyuan, et al.
Veröffentlicht: (2025) -
Zero-shot Vision-Language Reranking for Cross-View Geolocalization
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026) -
Zero-shot detection of buildings in mobile LiDAR using Language Vision Model
von: Goo, June Moh, et al.
Veröffentlicht: (2024) -
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
von: Guo, Grace, et al.
Veröffentlicht: (2024) -
Zeus: Zero-shot LLM Instruction for Union Segmentation in Multimodal Medical Imaging
von: Dai, Siyuan, et al.
Veröffentlicht: (2025)