Semantic Compositions Enhance Vision-Language Contrastive Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Aladago, Maxwell, Torresani, Lorenzo, Vosoughi, Soroush |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Open-world Instance Segmentation: Top-down Learning with Bottom-up Supervision
por: Kalluri, Tarun, et al.
Publicado: (2023)
por: Kalluri, Tarun, et al.
Publicado: (2023)
Compositional Entailment Learning for Hyperbolic Vision-Language Models
por: Pal, Avik, et al.
Publicado: (2024)
por: Pal, Avik, et al.
Publicado: (2024)
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
por: Diao, Xingjian, et al.
Publicado: (2025)
por: Diao, Xingjian, et al.
Publicado: (2025)
SMART: Semantic Matching Contrastive Learning for Partially View-Aligned Clustering
por: Peng, Liang, et al.
Publicado: (2025)
por: Peng, Liang, et al.
Publicado: (2025)
Frozen Vision Transformers for Dense Prediction on Small Datasets: A Case Study in Arrow Localization
por: Shepherd, Maxwell
Publicado: (2026)
por: Shepherd, Maxwell
Publicado: (2026)
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
por: Luo, Run, et al.
Publicado: (2025)
por: Luo, Run, et al.
Publicado: (2025)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
por: Lee, Jihoon, et al.
Publicado: (2025)
por: Lee, Jihoon, et al.
Publicado: (2025)
Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling
por: Hazimeh, Adam, et al.
Publicado: (2025)
por: Hazimeh, Adam, et al.
Publicado: (2025)
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
por: Lan, Zhibin, et al.
Publicado: (2025)
por: Lan, Zhibin, et al.
Publicado: (2025)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
por: Lavoie, Samuel, et al.
Publicado: (2024)
por: Lavoie, Samuel, et al.
Publicado: (2024)
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images
por: Wu, Shengguang, et al.
Publicado: (2025)
por: Wu, Shengguang, et al.
Publicado: (2025)
Candidate Pseudolabel Learning: Enhancing Vision-Language Models by Prompt Tuning with Unlabeled Data
por: Zhang, Jiahan, et al.
Publicado: (2024)
por: Zhang, Jiahan, et al.
Publicado: (2024)
Composition Vision-Language Understanding via Segment and Depth Anything Model
por: Huo, Mingxiao, et al.
Publicado: (2024)
por: Huo, Mingxiao, et al.
Publicado: (2024)
CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
por: Nguyen, Kiet A., et al.
Publicado: (2024)
por: Nguyen, Kiet A., et al.
Publicado: (2024)
Freeze the backbones: A Parameter-Efficient Contrastive Approach to Robust Medical Vision-Language Pre-training
por: Qin, Jiuming, et al.
Publicado: (2024)
por: Qin, Jiuming, et al.
Publicado: (2024)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
por: Li, Bangzheng, et al.
Publicado: (2025)
por: Li, Bangzheng, et al.
Publicado: (2025)
CLEFT: Language-Image Contrastive Learning with Efficient Large Language Model and Prompt Fine-Tuning
por: Du, Yuexi, et al.
Publicado: (2024)
por: Du, Yuexi, et al.
Publicado: (2024)
LayoutAgent: A Vision-Language Agent Guided Compositional Diffusion for Spatial Layout Planning
por: Fan, Zezhong, et al.
Publicado: (2025)
por: Fan, Zezhong, et al.
Publicado: (2025)
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
por: Fang, Yixiong, et al.
Publicado: (2024)
por: Fang, Yixiong, et al.
Publicado: (2024)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
por: Wan, David, et al.
Publicado: (2024)
por: Wan, David, et al.
Publicado: (2024)
VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis
por: Mehta, Manas, et al.
Publicado: (2025)
por: Mehta, Manas, et al.
Publicado: (2025)
Tree of Attributes Prompt Learning for Vision-Language Models
por: Ding, Tong, et al.
Publicado: (2024)
por: Ding, Tong, et al.
Publicado: (2024)
Efficient Few-Shot Learning in Remote Sensing: Fusing Vision and Vision-Language Models
por: Chua, Jia Yun, et al.
Publicado: (2025)
por: Chua, Jia Yun, et al.
Publicado: (2025)
Parallel In-context Learning for Large Vision Language Models
por: Yamaguchi, Shin'ya, et al.
Publicado: (2026)
por: Yamaguchi, Shin'ya, et al.
Publicado: (2026)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
por: Trager, Matthew, et al.
Publicado: (2023)
por: Trager, Matthew, et al.
Publicado: (2023)
Non-negative Contrastive Learning
por: Wang, Yifei, et al.
Publicado: (2024)
por: Wang, Yifei, et al.
Publicado: (2024)
AAPL: Adding Attributes to Prompt Learning for Vision-Language Models
por: Kim, Gahyeon, et al.
Publicado: (2024)
por: Kim, Gahyeon, et al.
Publicado: (2024)
Lightweight Unsupervised Federated Learning with Pretrained Vision Language Model
por: Yan, Hao, et al.
Publicado: (2024)
por: Yan, Hao, et al.
Publicado: (2024)
Vision-Language Models Provide Promptable Representations for Reinforcement Learning
por: Chen, William, et al.
Publicado: (2024)
por: Chen, William, et al.
Publicado: (2024)
Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
por: Kim, Gahyeon, et al.
Publicado: (2025)
por: Kim, Gahyeon, et al.
Publicado: (2025)
On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''
por: Bakker, Hua Chang, et al.
Publicado: (2025)
por: Bakker, Hua Chang, et al.
Publicado: (2025)
A Retrospect to Multi-prompt Learning across Vision and Language
por: Chen, Ziliang, et al.
Publicado: (2025)
por: Chen, Ziliang, et al.
Publicado: (2025)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
por: Pach, Mateusz, et al.
Publicado: (2025)
por: Pach, Mateusz, et al.
Publicado: (2025)
MuseCL: Predicting Urban Socioeconomic Indicators via Multi-Semantic Contrastive Learning
por: Yong, Xixian, et al.
Publicado: (2024)
por: Yong, Xixian, et al.
Publicado: (2024)
Adaptive Multi-head Contrastive Learning
por: Wang, Lei, et al.
Publicado: (2023)
por: Wang, Lei, et al.
Publicado: (2023)
Efficient Contrastive Decoding with Probabilistic Hallucination Detection - Mitigating Hallucinations in Large Vision Language Models -
por: Fieback, Laura, et al.
Publicado: (2025)
por: Fieback, Laura, et al.
Publicado: (2025)
Continual Learning in Vision-Language Models via Aligned Model Merging
por: Sokar, Ghada, et al.
Publicado: (2025)
por: Sokar, Ghada, et al.
Publicado: (2025)
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
por: Hinojosa, Carlos, et al.
Publicado: (2026)
por: Hinojosa, Carlos, et al.
Publicado: (2026)
I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation
por: Sassoon, Jordan, et al.
Publicado: (2025)
por: Sassoon, Jordan, et al.
Publicado: (2025)
Scaling Semantic Categories: Investigating the Impact on Vision Transformer Labeling Performance
por: Lamelas, Anthony, et al.
Publicado: (2025)
por: Lamelas, Anthony, et al.
Publicado: (2025)
Ejemplares similares
-
Open-world Instance Segmentation: Top-down Learning with Bottom-up Supervision
por: Kalluri, Tarun, et al.
Publicado: (2023) -
Compositional Entailment Learning for Hyperbolic Vision-Language Models
por: Pal, Avik, et al.
Publicado: (2024) -
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
por: Diao, Xingjian, et al.
Publicado: (2025) -
SMART: Semantic Matching Contrastive Learning for Partially View-Aligned Clustering
por: Peng, Liang, et al.
Publicado: (2025) -
Frozen Vision Transformers for Dense Prediction on Small Datasets: A Case Study in Arrow Localization
por: Shepherd, Maxwell
Publicado: (2026)