Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Berasi, Davide, Farina, Matteo, Mancini, Massimiliano, Ricci, Elisa, Strisciuglio, Nicola |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
von: Farina, Matteo, et al.
Veröffentlicht: (2025)
von: Farina, Matteo, et al.
Veröffentlicht: (2025)
Linear Model Merging Unlocks Simple and Scalable Multimodal Data Mixture Optimization
von: Berasi, Davide, et al.
Veröffentlicht: (2026)
von: Berasi, Davide, et al.
Veröffentlicht: (2026)
Frustratingly Easy Test-Time Adaptation of Vision-Language Models
von: Farina, Matteo, et al.
Veröffentlicht: (2024)
von: Farina, Matteo, et al.
Veröffentlicht: (2024)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
von: Guimard, Quentin, et al.
Veröffentlicht: (2026)
von: Guimard, Quentin, et al.
Veröffentlicht: (2026)
MULTIFLOW: Shifting Towards Task-Agnostic Vision-Language Pruning
von: Farina, Matteo, et al.
Veröffentlicht: (2024)
von: Farina, Matteo, et al.
Veröffentlicht: (2024)
Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers
von: Guimard, Quentin, et al.
Veröffentlicht: (2025)
von: Guimard, Quentin, et al.
Veröffentlicht: (2025)
Regressing Transformers for Data-efficient Visual Place Recognition
von: Leyva-Vallina, María, et al.
Veröffentlicht: (2024)
von: Leyva-Vallina, María, et al.
Veröffentlicht: (2024)
Large Multimodal Models as General In-Context Classifiers
von: Garosi, Marco, et al.
Veröffentlicht: (2026)
von: Garosi, Marco, et al.
Veröffentlicht: (2026)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
von: Caldarella, Simone, et al.
Veröffentlicht: (2024)
von: Caldarella, Simone, et al.
Veröffentlicht: (2024)
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
von: Vaish, Puru, et al.
Veröffentlicht: (2024)
von: Vaish, Puru, et al.
Veröffentlicht: (2024)
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
von: Das, Deepayan, et al.
Veröffentlicht: (2024)
von: Das, Deepayan, et al.
Veröffentlicht: (2024)
Specificity-aware reinforcement learning for fine-grained open-world classification
von: Angheben, Samuele, et al.
Veröffentlicht: (2026)
von: Angheben, Samuele, et al.
Veröffentlicht: (2026)
Debiasing Vison-Language Models with Text-Only Training
von: Yang, Yunfan, et al.
Veröffentlicht: (2024)
von: Yang, Yunfan, et al.
Veröffentlicht: (2024)
Concept-Aware Batch Sampling Improves Language-Image Pretraining
von: Ghosh, Adhiraj, et al.
Veröffentlicht: (2025)
von: Ghosh, Adhiraj, et al.
Veröffentlicht: (2025)
Can Text-to-Video Generation help Video-Language Alignment?
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
Mostly Text, Smart Visuals: Asymmetric Text-Visual Pruning for Large Vision-Language Models
von: Li, Sijie, et al.
Veröffentlicht: (2026)
von: Li, Sijie, et al.
Veröffentlicht: (2026)
Compositional Caching for Training-free Open-vocabulary Attribute Detection
von: Garosi, Marco, et al.
Veröffentlicht: (2025)
von: Garosi, Marco, et al.
Veröffentlicht: (2025)
Harnessing Large Language Models for Training-free Video Anomaly Detection
von: Zanella, Luca, et al.
Veröffentlicht: (2024)
von: Zanella, Luca, et al.
Veröffentlicht: (2024)
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
von: Das, Deepayan, et al.
Veröffentlicht: (2025)
von: Das, Deepayan, et al.
Veröffentlicht: (2025)
Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models
von: Huang, Xin, et al.
Veröffentlicht: (2025)
von: Huang, Xin, et al.
Veröffentlicht: (2025)
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
von: De Min, Thomas, et al.
Veröffentlicht: (2026)
von: De Min, Thomas, et al.
Veröffentlicht: (2026)
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
von: Uselis, Arnas, et al.
Veröffentlicht: (2026)
von: Uselis, Arnas, et al.
Veröffentlicht: (2026)
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
Vision-by-Language for Training-Free Compositional Image Retrieval
von: Karthik, Shyamgopal, et al.
Veröffentlicht: (2023)
von: Karthik, Shyamgopal, et al.
Veröffentlicht: (2023)
Visual Text Meets Low-level Vision: A Comprehensive Survey on Visual Text Processing
von: Shu, Yan, et al.
Veröffentlicht: (2024)
von: Shu, Yan, et al.
Veröffentlicht: (2024)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
von: Zhu, Lianghui, et al.
Veröffentlicht: (2024)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2024)
Attribute-based Visual Reprogramming for Vision-Language Models
von: Cai, Chengyi, et al.
Veröffentlicht: (2025)
von: Cai, Chengyi, et al.
Veröffentlicht: (2025)
MMRL: Multi-Modal Representation Learning for Vision-Language Models
von: Guo, Yuncheng, et al.
Veröffentlicht: (2025)
von: Guo, Yuncheng, et al.
Veröffentlicht: (2025)
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
von: Miranda, Imanol, et al.
Veröffentlicht: (2024)
von: Miranda, Imanol, et al.
Veröffentlicht: (2024)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
von: Huang, Chen, et al.
Veröffentlicht: (2026)
von: Huang, Chen, et al.
Veröffentlicht: (2026)
Towards Interpreting Visual Information Processing in Vision-Language Models
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
PHyCLIP: $\ell_1$-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
von: Yoshikawa, Daiki, et al.
Veröffentlicht: (2025)
von: Yoshikawa, Daiki, et al.
Veröffentlicht: (2025)
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
von: Mamaghan, Amir Mohammad Karimi, et al.
Veröffentlicht: (2024)
von: Mamaghan, Amir Mohammad Karimi, et al.
Veröffentlicht: (2024)
A Survey on the Robustness of Computer Vision Models against Common Corruptions
von: Wang, Shunxin, et al.
Veröffentlicht: (2023)
von: Wang, Shunxin, et al.
Veröffentlicht: (2023)
Multilingual Diversity Improves Vision-Language Representations
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
EasyARC: Evaluating Vision Language Models on True Visual Reasoning
von: Unsal, Mert, et al.
Veröffentlicht: (2025)
von: Unsal, Mert, et al.
Veröffentlicht: (2025)
VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
von: Hou, Haowen, et al.
Veröffentlicht: (2024)
von: Hou, Haowen, et al.
Veröffentlicht: (2024)
LT-Soups: Bridging Head and Tail Classes via Subsampled Model Soups
von: Aminbeidokhti, Masih, et al.
Veröffentlicht: (2025)
von: Aminbeidokhti, Masih, et al.
Veröffentlicht: (2025)
Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
von: Wang, Ruiyu, et al.
Veröffentlicht: (2025)
von: Wang, Ruiyu, et al.
Veröffentlicht: (2025)
Continual Learning of Conjugated Visual Representations through Higher-order Motion Flows
von: Marullo, Simone, et al.
Veröffentlicht: (2024)
von: Marullo, Simone, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
von: Farina, Matteo, et al.
Veröffentlicht: (2025) -
Linear Model Merging Unlocks Simple and Scalable Multimodal Data Mixture Optimization
von: Berasi, Davide, et al.
Veröffentlicht: (2026) -
Frustratingly Easy Test-Time Adaptation of Vision-Language Models
von: Farina, Matteo, et al.
Veröffentlicht: (2024) -
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
von: Guimard, Quentin, et al.
Veröffentlicht: (2026) -
MULTIFLOW: Shifting Towards Task-Agnostic Vision-Language Pruning
von: Farina, Matteo, et al.
Veröffentlicht: (2024)