Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
Fuente:
arXiv
Salvato in:
| Autori principali: | Ghate, Kshitish, Slaughter, Isaac, Wilson, Kyra, Diab, Mona, Caliskan, Aylin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes
di: Ghate, Kshitish, et al.
Pubblicazione: (2025)
di: Ghate, Kshitish, et al.
Pubblicazione: (2025)
Personal Information Parroting in Language Models
di: Subramani, Nishant, et al.
Pubblicazione: (2026)
di: Subramani, Nishant, et al.
Pubblicazione: (2026)
Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval
di: Wilson, Kyra, et al.
Pubblicazione: (2024)
di: Wilson, Kyra, et al.
Pubblicazione: (2024)
Bias Amplification in Stable Diffusion's Representation of Stigma Through Skin Tones and Their Homogeneity
di: Wilson, Kyra, et al.
Pubblicazione: (2025)
di: Wilson, Kyra, et al.
Pubblicazione: (2025)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
di: Chopra, Harshita, et al.
Pubblicazione: (2026)
di: Chopra, Harshita, et al.
Pubblicazione: (2026)
Generative Value Conflicts Reveal LLM Priorities
di: Liu, Andy, et al.
Pubblicazione: (2025)
di: Liu, Andy, et al.
Pubblicazione: (2025)
Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models
di: Veerendranath, Vishruth, et al.
Pubblicazione: (2024)
di: Veerendranath, Vishruth, et al.
Pubblicazione: (2024)
ChatGPT Perpetuates Gender Bias in Machine Translation and Ignores Non-Gendered Pronouns: Findings across Bengali and Five other Low-Resource Languages
di: Ghosh, Sourojit, et al.
Pubblicazione: (2023)
di: Ghosh, Sourojit, et al.
Pubblicazione: (2023)
Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
di: Gueorguieva, Anna-Maria, et al.
Pubblicazione: (2025)
di: Gueorguieva, Anna-Maria, et al.
Pubblicazione: (2025)
No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy
di: Wilson, Kyra, et al.
Pubblicazione: (2025)
di: Wilson, Kyra, et al.
Pubblicazione: (2025)
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
di: Ghate, Kshitish, et al.
Pubblicazione: (2025)
di: Ghate, Kshitish, et al.
Pubblicazione: (2025)
Deep Reasoning in General Purpose Agents via Structured Meta-Cognition
di: Light, Dean, et al.
Pubblicazione: (2026)
di: Light, Dean, et al.
Pubblicazione: (2026)
A Taxonomy of Stereotype Content in Large Language Models
di: Nicolas, Gandalf, et al.
Pubblicazione: (2024)
di: Nicolas, Gandalf, et al.
Pubblicazione: (2024)
Finetune-Informed Pretraining Boosts Downstream Performance
di: Faysal, Atik, et al.
Pubblicazione: (2026)
di: Faysal, Atik, et al.
Pubblicazione: (2026)
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning
di: Dash, Saloni, et al.
Pubblicazione: (2025)
di: Dash, Saloni, et al.
Pubblicazione: (2025)
Renaissance: Investigating the Pretraining of Vision-Language Encoders
di: Fields, Clayton, et al.
Pubblicazione: (2024)
di: Fields, Clayton, et al.
Pubblicazione: (2024)
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
di: Subramani, Nishant, et al.
Pubblicazione: (2025)
di: Subramani, Nishant, et al.
Pubblicazione: (2025)
Emotion Classification in Low and Moderate Resource Languages
di: Tafreshi, Shabnam, et al.
Pubblicazione: (2024)
di: Tafreshi, Shabnam, et al.
Pubblicazione: (2024)
Do Generative AI Models Output Harm while Representing Non-Western Cultures: Evidence from A Community-Centered Approach
di: Ghosh, Sourojit, et al.
Pubblicazione: (2024)
di: Ghosh, Sourojit, et al.
Pubblicazione: (2024)
REALM: A Dataset of Real-World LLM Use Cases
di: Cheng, Jingwen, et al.
Pubblicazione: (2025)
di: Cheng, Jingwen, et al.
Pubblicazione: (2025)
StressRoBERTa: Cross-Condition Transfer Learning from Depression, Anxiety, and PTSD to Stress Detection
di: Alqahtani, Amal, et al.
Pubblicazione: (2025)
di: Alqahtani, Amal, et al.
Pubblicazione: (2025)
Can Large Language Models Infer Causation from Correlation?
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
Vision-Language Model Selection and Reuse for Downstream Adaptation
di: Tan, Hao-Zhe, et al.
Pubblicazione: (2025)
di: Tan, Hao-Zhe, et al.
Pubblicazione: (2025)
VEAT Quantifies Implicit Associations in Text-to-Video Generator Sora and Reveals Challenges in Bias Mitigation
di: Sun, Yongxu, et al.
Pubblicazione: (2026)
di: Sun, Yongxu, et al.
Pubblicazione: (2026)
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
di: Wedgwood, James, et al.
Pubblicazione: (2026)
di: Wedgwood, James, et al.
Pubblicazione: (2026)
Scaling Laws for Predicting Downstream Performance in LLMs
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
di: Muhamed, Aashiq, et al.
Pubblicazione: (2024)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2024)
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
di: Touchent, Rian, et al.
Pubblicazione: (2026)
di: Touchent, Rian, et al.
Pubblicazione: (2026)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
di: Li, Miaomiao, et al.
Pubblicazione: (2025)
di: Li, Miaomiao, et al.
Pubblicazione: (2025)
How Early Adopters Used Generative AI Worldwide: Variation by Country Income and Language
di: Daepp, Madeleine I. G., et al.
Pubblicazione: (2026)
di: Daepp, Madeleine I. G., et al.
Pubblicazione: (2026)
CoRAG: Collaborative Retrieval-Augmented Generation
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
Harnessing the Intrinsic Knowledge of Pretrained Language Models for Challenging Text Classification Settings
di: Gao, Lingyu
Pubblicazione: (2024)
di: Gao, Lingyu
Pubblicazione: (2024)
When Do LLM Preferences Predict Downstream Behavior?
di: Slama, Katarina, et al.
Pubblicazione: (2026)
di: Slama, Katarina, et al.
Pubblicazione: (2026)
Quantifying the Importance of Data Alignment in Downstream Model Performance
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
Downstream Transfer Attack: Adversarial Attacks on Downstream Models with Pre-trained Vision Transformers
di: Zheng, Weijie, et al.
Pubblicazione: (2024)
di: Zheng, Weijie, et al.
Pubblicazione: (2024)
Locating Demographic Bias at the Attention-Head Level in CLIP's Vision Encoder
di: Yasser, Alaa, et al.
Pubblicazione: (2026)
di: Yasser, Alaa, et al.
Pubblicazione: (2026)
Piano Transcription by Hierarchical Language Modeling with Pretrained Roll-based Encoders
di: Li, Dichucheng, et al.
Pubblicazione: (2025)
di: Li, Dichucheng, et al.
Pubblicazione: (2025)
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
di: An, Na Min, et al.
Pubblicazione: (2025)
di: An, Na Min, et al.
Pubblicazione: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
di: Lavoie, Samuel, et al.
Pubblicazione: (2024)
di: Lavoie, Samuel, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes
di: Ghate, Kshitish, et al.
Pubblicazione: (2025) -
Personal Information Parroting in Language Models
di: Subramani, Nishant, et al.
Pubblicazione: (2026) -
Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval
di: Wilson, Kyra, et al.
Pubblicazione: (2024) -
Bias Amplification in Stable Diffusion's Representation of Stigma Through Skin Tones and Their Homogeneity
di: Wilson, Kyra, et al.
Pubblicazione: (2025) -
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
di: Chopra, Harshita, et al.
Pubblicazione: (2026)