CountCLIP -- [Re] Teaching CLIP to Count to Ten
Fuente:
arXiv
Salvato in:
| Autori principali: | Mestha, Harshvardhan, Agrawal, Tejas, Bania, Karan, V, Shreyas, Bhisikar, Yash |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
STREAM: A Universal State-Space Model for Sparse Geometric Data
di: Schöne, Mark, et al.
Pubblicazione: (2024)
di: Schöne, Mark, et al.
Pubblicazione: (2024)
Multi-Turn Human-LLM Interaction Through the Lens of a Two-Way Intelligibility Protocol
di: Mestha, Harshvardhan, et al.
Pubblicazione: (2024)
di: Mestha, Harshvardhan, et al.
Pubblicazione: (2024)
Reproducibility Study of CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image Classification
di: Shah, Manan, et al.
Pubblicazione: (2024)
di: Shah, Manan, et al.
Pubblicazione: (2024)
DiffCLIP: Differential Attention Meets CLIP
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
di: Lai, Zhengfeng, et al.
Pubblicazione: (2023)
di: Lai, Zhengfeng, et al.
Pubblicazione: (2023)
CLIP Can Understand Depth
di: Kim, Sohee, et al.
Pubblicazione: (2024)
di: Kim, Sohee, et al.
Pubblicazione: (2024)
ECOR: Explainable CLIP for Object Recognition
di: Rasekh, Ali, et al.
Pubblicazione: (2024)
di: Rasekh, Ali, et al.
Pubblicazione: (2024)
Interpreting CLIP with Hierarchical Sparse Autoencoders
di: Zaigrajew, Vladimir, et al.
Pubblicazione: (2025)
di: Zaigrajew, Vladimir, et al.
Pubblicazione: (2025)
Implicit Inversion turns CLIP into a Decoder
di: D'Orazio, Antonio, et al.
Pubblicazione: (2025)
di: D'Orazio, Antonio, et al.
Pubblicazione: (2025)
Steering CLIP's vision transformer with sparse autoencoders
di: Joseph, Sonia, et al.
Pubblicazione: (2025)
di: Joseph, Sonia, et al.
Pubblicazione: (2025)
IDEA: Image Description Enhanced CLIP-Adapter
di: Ye, Zhipeng, et al.
Pubblicazione: (2025)
di: Ye, Zhipeng, et al.
Pubblicazione: (2025)
RankCLIP: Ranking-Consistent Language-Image Pretraining
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
di: Yang, Tianyu, et al.
Pubblicazione: (2024)
di: Yang, Tianyu, et al.
Pubblicazione: (2024)
WalkCLIP: Multimodal Learning for Urban Walkability Prediction
di: Xiang, Shilong, et al.
Pubblicazione: (2025)
di: Xiang, Shilong, et al.
Pubblicazione: (2025)
Learning Generalizable Prompt for CLIP with Class Similarity Knowledge
di: Jung, Sehun, et al.
Pubblicazione: (2025)
di: Jung, Sehun, et al.
Pubblicazione: (2025)
Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings
di: Kim, Bumjun, et al.
Pubblicazione: (2026)
di: Kim, Bumjun, et al.
Pubblicazione: (2026)
Captured by Captions: On Memorization and its Mitigation in CLIP Models
di: Wang, Wenhao, et al.
Pubblicazione: (2025)
di: Wang, Wenhao, et al.
Pubblicazione: (2025)
AtGCN: A Graph Convolutional Network For Ataxic Gait Detection
di: Bania, Karan, et al.
Pubblicazione: (2024)
di: Bania, Karan, et al.
Pubblicazione: (2024)
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
di: Kim, Hyunjae, et al.
Pubblicazione: (2024)
di: Kim, Hyunjae, et al.
Pubblicazione: (2024)
On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''
di: Bakker, Hua Chang, et al.
Pubblicazione: (2025)
di: Bakker, Hua Chang, et al.
Pubblicazione: (2025)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
di: Crabbé, Jonathan, et al.
Pubblicazione: (2023)
di: Crabbé, Jonathan, et al.
Pubblicazione: (2023)
A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation
di: Wang, Zhengbo, et al.
Pubblicazione: (2024)
di: Wang, Zhengbo, et al.
Pubblicazione: (2024)
Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone Ensembling
di: Rodriguez-Opazo, Cristian, et al.
Pubblicazione: (2024)
di: Rodriguez-Opazo, Cristian, et al.
Pubblicazione: (2024)
CLIP with Generative Latent Replay: a Strong Baseline for Incremental Learning
di: Frascaroli, Emanuele, et al.
Pubblicazione: (2024)
di: Frascaroli, Emanuele, et al.
Pubblicazione: (2024)
Robustness in Both Domains: CLIP Needs a Robust Text Encoder
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
CLIP-Inspector: Model-Level Backdoor Detection for Prompt-Tuned CLIP via OOD Trigger Inversion
di: Jindal, Akshit, et al.
Pubblicazione: (2026)
di: Jindal, Akshit, et al.
Pubblicazione: (2026)
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
di: Mistretta, Marco, et al.
Pubblicazione: (2025)
di: Mistretta, Marco, et al.
Pubblicazione: (2025)
COOkeD: Ensemble-based OOD detection in the era of zero-shot CLIP
di: Humblot-Renaux, Galadrielle, et al.
Pubblicazione: (2025)
di: Humblot-Renaux, Galadrielle, et al.
Pubblicazione: (2025)
GoldiCLIP: The Goldilocks Approach for Balancing Explicit Supervision for Language-Image Pretraining
di: Mohan, Deen Dayal, et al.
Pubblicazione: (2026)
di: Mohan, Deen Dayal, et al.
Pubblicazione: (2026)
ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection
di: Ma, Ke, et al.
Pubblicazione: (2025)
di: Ma, Ke, et al.
Pubblicazione: (2025)
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
di: Metzen, Jan Hendrik, et al.
Pubblicazione: (2023)
di: Metzen, Jan Hendrik, et al.
Pubblicazione: (2023)
Enhancing CLIP Conceptual Embedding through Knowledge Distillation
di: Kao, Kuei-Chun
Pubblicazione: (2024)
di: Kao, Kuei-Chun
Pubblicazione: (2024)
MoDE: CLIP Data Experts via Clustering
di: Ma, Jiawei, et al.
Pubblicazione: (2024)
di: Ma, Jiawei, et al.
Pubblicazione: (2024)
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
di: An, Bang, et al.
Pubblicazione: (2023)
di: An, Bang, et al.
Pubblicazione: (2023)
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2023)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2023)
SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIP
di: Timmermann, Christoph, et al.
Pubblicazione: (2025)
di: Timmermann, Christoph, et al.
Pubblicazione: (2025)
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
di: Xie, Chunyu, et al.
Pubblicazione: (2025)
di: Xie, Chunyu, et al.
Pubblicazione: (2025)
MobileCLIP2: Improving Multi-Modal Reinforced Training
di: Faghri, Fartash, et al.
Pubblicazione: (2025)
di: Faghri, Fartash, et al.
Pubblicazione: (2025)
Documenti analoghi
-
STREAM: A Universal State-Space Model for Sparse Geometric Data
di: Schöne, Mark, et al.
Pubblicazione: (2024) -
Multi-Turn Human-LLM Interaction Through the Lens of a Two-Way Intelligibility Protocol
di: Mestha, Harshvardhan, et al.
Pubblicazione: (2024) -
Reproducibility Study of CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image Classification
di: Shah, Manan, et al.
Pubblicazione: (2024) -
DiffCLIP: Differential Attention Meets CLIP
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025) -
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)