Who's in and who's out? A case study of multimodal CLIP-filtering in DataComp
Fuente:
arXiv
Salvato in:
| Autori principali: | Hong, Rachel, Agnew, William, Kohno, Tadayoshi, Morgenstern, Jamie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
di: Hong, Rachel, et al.
Pubblicazione: (2025)
di: Hong, Rachel, et al.
Pubblicazione: (2025)
Demystifying CLIP Data
di: Xu, Hu, et al.
Pubblicazione: (2023)
di: Xu, Hu, et al.
Pubblicazione: (2023)
Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking
di: Nemecek, Alexander, et al.
Pubblicazione: (2026)
di: Nemecek, Alexander, et al.
Pubblicazione: (2026)
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
di: Zhang, Sheng, et al.
Pubblicazione: (2023)
di: Zhang, Sheng, et al.
Pubblicazione: (2023)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
di: Chew, Oscar, et al.
Pubblicazione: (2026)
di: Chew, Oscar, et al.
Pubblicazione: (2026)
Will we run out of data? Limits of LLM scaling based on human-generated data
di: Villalobos, Pablo, et al.
Pubblicazione: (2022)
di: Villalobos, Pablo, et al.
Pubblicazione: (2022)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
di: Park, Junsung, et al.
Pubblicazione: (2025)
di: Park, Junsung, et al.
Pubblicazione: (2025)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
di: Huang, Yuchen, et al.
Pubblicazione: (2025)
di: Huang, Yuchen, et al.
Pubblicazione: (2025)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
di: Lucy, Li, et al.
Pubblicazione: (2026)
di: Lucy, Li, et al.
Pubblicazione: (2026)
Stable Signer: Hierarchical Sign Language Generative Model
di: Fang, Sen, et al.
Pubblicazione: (2025)
di: Fang, Sen, et al.
Pubblicazione: (2025)
Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
di: Fraser, Kathleen C., et al.
Pubblicazione: (2024)
di: Fraser, Kathleen C., et al.
Pubblicazione: (2024)
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
di: Jha, Akshita, et al.
Pubblicazione: (2024)
di: Jha, Akshita, et al.
Pubblicazione: (2024)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
di: Sathe, Ashutosh, et al.
Pubblicazione: (2024)
di: Sathe, Ashutosh, et al.
Pubblicazione: (2024)
The Use of Multimodal Large Language Models to Detect Objects from Thermal Images: Transportation Applications
di: Ashqar, Huthaifa I., et al.
Pubblicazione: (2024)
di: Ashqar, Huthaifa I., et al.
Pubblicazione: (2024)
FigSIM: A Dataset for Fine-grained Suicide Severity and Figurative Language in Suicide Memes
di: Chen, Liuliu, et al.
Pubblicazione: (2026)
di: Chen, Liuliu, et al.
Pubblicazione: (2026)
ERIT Lightweight Multimodal Dataset for Elderly Emotion Recognition and Multimodal Fusion Evaluation
di: Frieske, Rita, et al.
Pubblicazione: (2024)
di: Frieske, Rita, et al.
Pubblicazione: (2024)
Large Language Models and Provenance Metadata for Determining the Relevance of Images and Videos in News Stories
di: Peterka, Tomas, et al.
Pubblicazione: (2025)
di: Peterka, Tomas, et al.
Pubblicazione: (2025)
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
di: Nazi, Zabir Al, et al.
Pubblicazione: (2025)
di: Nazi, Zabir Al, et al.
Pubblicazione: (2025)
Using LLMs as prompt modifier to avoid biases in AI image generators
di: Peinl, René
Pubblicazione: (2025)
di: Peinl, René
Pubblicazione: (2025)
Restoring Ancient Ideograph: A Multimodal Multitask Neural Network Approach
di: Duan, Siyu, et al.
Pubblicazione: (2024)
di: Duan, Siyu, et al.
Pubblicazione: (2024)
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
di: Si, Chenglei, et al.
Pubblicazione: (2024)
di: Si, Chenglei, et al.
Pubblicazione: (2024)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
di: Jin, Yiqiao, et al.
Pubblicazione: (2024)
di: Jin, Yiqiao, et al.
Pubblicazione: (2024)
A Longitudinal Analysis of Racial and Gender Bias in New York Times and Fox News Images and Articles
di: Ibrahim, Hazem, et al.
Pubblicazione: (2024)
di: Ibrahim, Hazem, et al.
Pubblicazione: (2024)
SideSeeing: A multimodal dataset and collection of tools for sidewalk assessment
di: Damaceno, R. J. P., et al.
Pubblicazione: (2024)
di: Damaceno, R. J. P., et al.
Pubblicazione: (2024)
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task
di: Asgarov, Ali, et al.
Pubblicazione: (2024)
di: Asgarov, Ali, et al.
Pubblicazione: (2024)
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
di: Xie, Jingyou, et al.
Pubblicazione: (2024)
di: Xie, Jingyou, et al.
Pubblicazione: (2024)
Closing the Gap: Data-Centric Fine-Tuning of Vision Language Models for the Standardized Exam Questions
di: Sert, Egemen, et al.
Pubblicazione: (2025)
di: Sert, Egemen, et al.
Pubblicazione: (2025)
D3G: Diverse Demographic Data Generation Increases Zero-Shot Image Classification Accuracy within Multimodal Models
di: Hickmon, Javon
Pubblicazione: (2025)
di: Hickmon, Javon
Pubblicazione: (2025)
Detecting Visual Triggers in Cannabis Imagery: A CLIP-Based Multi-Labeling Framework with Local-Global Aggregation
di: Lu, Linqi, et al.
Pubblicazione: (2024)
di: Lu, Linqi, et al.
Pubblicazione: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
di: Wang, Jiapeng, et al.
Pubblicazione: (2024)
di: Wang, Jiapeng, et al.
Pubblicazione: (2024)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
Meta CLIP 2: A Worldwide Scaling Recipe
di: Chuang, Yung-Sung, et al.
Pubblicazione: (2025)
di: Chuang, Yung-Sung, et al.
Pubblicazione: (2025)
Generalizable Prompt Learning of CLIP: A Brief Overview
di: Cui, Fangming, et al.
Pubblicazione: (2025)
di: Cui, Fangming, et al.
Pubblicazione: (2025)
Architecture inside the mirage: evaluating generative image models on architectural style, elements, and typologies
di: Magrill, Jamie, et al.
Pubblicazione: (2026)
di: Magrill, Jamie, et al.
Pubblicazione: (2026)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
di: Yan, Zehong, et al.
Pubblicazione: (2025)
di: Yan, Zehong, et al.
Pubblicazione: (2025)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
di: Gao, Peng, et al.
Pubblicazione: (2021)
di: Gao, Peng, et al.
Pubblicazione: (2021)
TiC-CLIP: Continual Training of CLIP Models
di: Garg, Saurabh, et al.
Pubblicazione: (2023)
di: Garg, Saurabh, et al.
Pubblicazione: (2023)
Scaling medical imaging report generation with multimodal reinforcement learning
di: Liu, Qianchu, et al.
Pubblicazione: (2026)
di: Liu, Qianchu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
di: Hong, Rachel, et al.
Pubblicazione: (2025) -
Demystifying CLIP Data
di: Xu, Hu, et al.
Pubblicazione: (2023) -
Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking
di: Nemecek, Alexander, et al.
Pubblicazione: (2026) -
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
di: Zhang, Sheng, et al.
Pubblicazione: (2023) -
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)