Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
Fuente:
arXiv
Salvato in:
| Autori principali: | Mayilvahanan, Prasanna, Wiedemer, Thaddäus, Rusak, Evgenia, Bethge, Matthias, Brendel, Wieland |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
In Search of Forgotten Domain Generalization
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2024)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2024)
MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
di: Zeller, Jana, et al.
Pubblicazione: (2026)
di: Zeller, Jana, et al.
Pubblicazione: (2026)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
VGGSounder: Audio-Visual Evaluations for Foundation Models
di: Zverev, Daniil, et al.
Pubblicazione: (2025)
di: Zverev, Daniil, et al.
Pubblicazione: (2025)
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
Low-Pass Filtering Improves Behavioral Alignment of Vision Models
di: Wolff, Max, et al.
Pubblicazione: (2026)
di: Wolff, Max, et al.
Pubblicazione: (2026)
Effective pruning of web-scale datasets based on complexity of concept clusters
di: Abbas, Amro, et al.
Pubblicazione: (2024)
di: Abbas, Amro, et al.
Pubblicazione: (2024)
InfoNCE: Identifying the Gap Between Theory and Practice
di: Rusak, Evgenia, et al.
Pubblicazione: (2024)
di: Rusak, Evgenia, et al.
Pubblicazione: (2024)
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
di: Cai, Yuliang, et al.
Pubblicazione: (2025)
di: Cai, Yuliang, et al.
Pubblicazione: (2025)
Provable Compositional Generalization for Object-Centric Learning
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2023)
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2023)
Scale Alone Does not Improve Mechanistic Interpretability in Vision Models
di: Zimmermann, Roland S., et al.
Pubblicazione: (2023)
di: Zimmermann, Roland S., et al.
Pubblicazione: (2023)
DocumentCLIP: Linking Figures and Main Body Text in Reflowed Documents
di: Liu, Fuxiao, et al.
Pubblicazione: (2023)
di: Liu, Fuxiao, et al.
Pubblicazione: (2023)
Video models are zero-shot learners and reasoners
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
Modeling Saliency Dataset Bias
di: Kümmerer, Matthias, et al.
Pubblicazione: (2025)
di: Kümmerer, Matthias, et al.
Pubblicazione: (2025)
Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality Assessment
di: Liao, Zhicheng, et al.
Pubblicazione: (2025)
di: Liao, Zhicheng, et al.
Pubblicazione: (2025)
Learning Generalizable Prompt for CLIP with Class Similarity Knowledge
di: Jung, Sehun, et al.
Pubblicazione: (2025)
di: Jung, Sehun, et al.
Pubblicazione: (2025)
Revealing the Underlying Patterns: Investigating Dataset Similarity, Performance, and Generalization
di: Achara, Akshit, et al.
Pubblicazione: (2023)
di: Achara, Akshit, et al.
Pubblicazione: (2023)
MINT: Memory-Infused Prompt Tuning at Test-time for CLIP
di: Yi, Jiaming, et al.
Pubblicazione: (2025)
di: Yi, Jiaming, et al.
Pubblicazione: (2025)
Are We Done with Object-Centric Learning?
di: Rubinstein, Alexander, et al.
Pubblicazione: (2025)
di: Rubinstein, Alexander, et al.
Pubblicazione: (2025)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
di: Lai, Zhengfeng, et al.
Pubblicazione: (2023)
di: Lai, Zhengfeng, et al.
Pubblicazione: (2023)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
di: Lewis, Martha, et al.
Pubblicazione: (2022)
di: Lewis, Martha, et al.
Pubblicazione: (2022)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
di: Gao, Bin-Bin, et al.
Pubblicazione: (2025)
di: Gao, Bin-Bin, et al.
Pubblicazione: (2025)
DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
di: Wang, Zhu, et al.
Pubblicazione: (2025)
di: Wang, Zhu, et al.
Pubblicazione: (2025)
Don't trust your eyes: on the (un)reliability of feature visualizations
di: Geirhos, Robert, et al.
Pubblicazione: (2023)
di: Geirhos, Robert, et al.
Pubblicazione: (2023)
Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic Segmentation
di: Lee, ByeongCheol, et al.
Pubblicazione: (2026)
di: Lee, ByeongCheol, et al.
Pubblicazione: (2026)
StableTTA: Improving Vision Model Performance by Training-free Test-Time Adaptation Methods
di: Li, Zheng, et al.
Pubblicazione: (2026)
di: Li, Zheng, et al.
Pubblicazione: (2026)
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
di: Madasu, Avinash, et al.
Pubblicazione: (2025)
di: Madasu, Avinash, et al.
Pubblicazione: (2025)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
AA-CLIP: Enhancing Zero-shot Anomaly Detection via Anomaly-Aware CLIP
di: Ma, Wenxin, et al.
Pubblicazione: (2025)
di: Ma, Wenxin, et al.
Pubblicazione: (2025)
CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic Decoding
di: Zhou, Qiongyi, et al.
Pubblicazione: (2024)
di: Zhou, Qiongyi, et al.
Pubblicazione: (2024)
Test-Time Adaptation with SaLIP: A Cascade of SAM and CLIP for Zero shot Medical Image Segmentation
di: Aleem, Sidra, et al.
Pubblicazione: (2024)
di: Aleem, Sidra, et al.
Pubblicazione: (2024)
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
di: Lee, Sangin, et al.
Pubblicazione: (2026)
di: Lee, Sangin, et al.
Pubblicazione: (2026)
CytoCLIP: Learning Cytoarchitectural Characteristics in Developing Human Brain Using Contrastive Language Image Pre-Training
di: Ta, Pralaypati, et al.
Pubblicazione: (2026)
di: Ta, Pralaypati, et al.
Pubblicazione: (2026)
Highly Compressed Tokenizer Can Generate Without Training
di: Beyer, L. Lao, et al.
Pubblicazione: (2025)
di: Beyer, L. Lao, et al.
Pubblicazione: (2025)
CRRG-CLIP: Automatic Generation of Chest Radiology Reports and Classification of Chest Radiographs
di: Xu, Jianfei, et al.
Pubblicazione: (2024)
di: Xu, Jianfei, et al.
Pubblicazione: (2024)
CLIP-DQA: Blindly Evaluating Dehazed Images from Global and Local Perspectives Using CLIP
di: Zeng, Yirui, et al.
Pubblicazione: (2025)
di: Zeng, Yirui, et al.
Pubblicazione: (2025)
FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
ComCLIP: Training-Free Compositional Image and Text Matching
di: Jiang, Kenan, et al.
Pubblicazione: (2022)
di: Jiang, Kenan, et al.
Pubblicazione: (2022)
Documenti analoghi
-
In Search of Forgotten Domain Generalization
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2024) -
MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
di: Zeller, Jana, et al.
Pubblicazione: (2026) -
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025) -
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025) -
VGGSounder: Audio-Visual Evaluations for Foundation Models
di: Zverev, Daniil, et al.
Pubblicazione: (2025)