The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ovalle, Anaelia, Pavasovic, Krunoslav Lehman, Martin, Louis, Zettlemoyer, Luke, Smith, Eric Michael, Chang, Kai-Wei, Williams, Adina, Sagun, Levent |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Differentiable Rank-Based Objective For Better Feature Learning
by: Pavasovic, Krunoslav Lehman, et al.
Published: (2025)
by: Pavasovic, Krunoslav Lehman, et al.
Published: (2025)
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
by: Ovalle, Anaelia, et al.
Published: (2025)
by: Ovalle, Anaelia, et al.
Published: (2025)
Classifier-Free Guidance: From High-Dimensional Analysis to Generalized Guidance Forms
by: Pavasovic, Krunoslav Lehman, et al.
Published: (2025)
by: Pavasovic, Krunoslav Lehman, et al.
Published: (2025)
Weisfeiler and Leman Go Measurement Modeling: Probing the Validity of the WL Test
by: Subramonian, Arjun, et al.
Published: (2023)
by: Subramonian, Arjun, et al.
Published: (2023)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
by: Ovalle, Anaelia, et al.
Published: (2023)
by: Ovalle, Anaelia, et al.
Published: (2023)
Chained Tuning Leads to Biased Forgetting
by: Ung, Megan, et al.
Published: (2024)
by: Ung, Megan, et al.
Published: (2024)
Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link Prediction
by: Subramonian, Arjun, et al.
Published: (2023)
by: Subramonian, Arjun, et al.
Published: (2023)
Reassessing the Validity of Spurious Correlations Benchmarks
by: Bell, Samuel J., et al.
Published: (2024)
by: Bell, Samuel J., et al.
Published: (2024)
Brittlebench: Quantifying LLM robustness via prompt sensitivity
by: Romanou, Angelika, et al.
Published: (2026)
by: Romanou, Angelika, et al.
Published: (2026)
Are Female Carpenters like Blue Bananas? A Corpus Investigation of Occupation Gender Typicality
by: Ju, Da, et al.
Published: (2024)
by: Ju, Da, et al.
Published: (2024)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
by: Yasunaga, Michihiro, et al.
Published: (2025)
by: Yasunaga, Michihiro, et al.
Published: (2025)
An Effective Theory of Bias Amplification
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
by: Shaib, Chantal, et al.
Published: (2025)
by: Shaib, Chantal, et al.
Published: (2025)
On the Role of Speech Data in Reducing Toxicity Detection Bias
by: Bell, Samuel J., et al.
Published: (2024)
by: Bell, Samuel J., et al.
Published: (2024)
Demystifying Prompts in Language Models via Perplexity Estimation
by: Gonen, Hila, et al.
Published: (2022)
by: Gonen, Hila, et al.
Published: (2022)
Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
by: Min, Sewon, et al.
Published: (2023)
by: Min, Sewon, et al.
Published: (2023)
Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language Models
by: Gonen, Hila, et al.
Published: (2024)
by: Gonen, Hila, et al.
Published: (2024)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
by: Chen, Tong, et al.
Published: (2025)
by: Chen, Tong, et al.
Published: (2025)
The Causal Influence of Grammatical Gender on Distributional Semantics
by: Stańczak, Karolina, et al.
Published: (2023)
by: Stańczak, Karolina, et al.
Published: (2023)
Small Divisor problems and $A_p$ weights with an application
by: Chanillo, Sagun
Published: (2024)
by: Chanillo, Sagun
Published: (2024)
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Evaluating Copyright Takedown Methods for Language Models
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
by: Kilian, Maciej, et al.
Published: (2024)
by: Kilian, Maciej, et al.
Published: (2024)
Comparing Hallucination Detection Metrics for Multilingual Generation
by: Kang, Haoqiang, et al.
Published: (2024)
by: Kang, Haoqiang, et al.
Published: (2024)
(Mis)Fitting: A Survey of Scaling Laws
by: Li, Margaret, et al.
Published: (2025)
by: Li, Margaret, et al.
Published: (2025)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
by: Kamath, Amita, et al.
Published: (2025)
by: Kamath, Amita, et al.
Published: (2025)
The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation
by: Mundt, Martin, et al.
Published: (2025)
by: Mundt, Martin, et al.
Published: (2025)
MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling
by: Limisiewicz, Tomasz, et al.
Published: (2024)
by: Limisiewicz, Tomasz, et al.
Published: (2024)
Breaking the Curse of Multilinguality with Cross-lingual Expert Language Models
by: Blevins, Terra, et al.
Published: (2024)
by: Blevins, Terra, et al.
Published: (2024)
Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
Chapter 1 Language and leadership
by: Lehman, Iga Maria
Published: (2024)
by: Lehman, Iga Maria
Published: (2024)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
by: Dorn, Rebecca, et al.
Published: (2024)
by: Dorn, Rebecca, et al.
Published: (2024)
GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Micro Language Models Enable Instant Responses
by: Cheng, Wen, et al.
Published: (2026)
by: Cheng, Wen, et al.
Published: (2026)
Sense and Sensitivity: Evaluating the simulation of social dynamics via Large Language Models
by: Ju, Da, et al.
Published: (2024)
by: Ju, Da, et al.
Published: (2024)
Sharp bounds on the Nusselt number in Rayleigh-Bénard convection and a bilinear estimate via Carleson measures
by: Chanillo, Sagun, et al.
Published: (2020)
by: Chanillo, Sagun, et al.
Published: (2020)
Slicing and Dicing: Configuring Optimal Mixtures of Experts
by: Li, Margaret, et al.
Published: (2026)
by: Li, Margaret, et al.
Published: (2026)
Gendered Predictors of Exclusive Breastfeeding Among Employed Mothers: An Ecological Multicenter Study
by: Bishayr Hassan Aljaffar, et al.
Published: (2025)
by: Bishayr Hassan Aljaffar, et al.
Published: (2025)
High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuning
by: Franzmeyer, Tim, et al.
Published: (2025)
by: Franzmeyer, Tim, et al.
Published: (2025)
Similar Items
-
A Differentiable Rank-Based Objective For Better Feature Learning
by: Pavasovic, Krunoslav Lehman, et al.
Published: (2025) -
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
by: Ovalle, Anaelia, et al.
Published: (2025) -
Classifier-Free Guidance: From High-Dimensional Analysis to Generalized Guidance Forms
by: Pavasovic, Krunoslav Lehman, et al.
Published: (2025) -
Weisfeiler and Leman Go Measurement Modeling: Probing the Validity of the WL Test
by: Subramonian, Arjun, et al.
Published: (2023) -
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
by: Ovalle, Anaelia, et al.
Published: (2023)