Breaking Down Bias: On The Limits of Generalizable Pruning Strategies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Sibo, Salinas, Alejandro, Henderson, Peter, Nyarko, Julian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What's in a Name? Auditing Large Language Models for Race and Gender Bias
von: Salinas, Alejandro, et al.
Veröffentlicht: (2024)
von: Salinas, Alejandro, et al.
Veröffentlicht: (2024)
Identifying Emerging Concepts in Large Corpora
von: Ma, Sibo, et al.
Veröffentlicht: (2025)
von: Ma, Sibo, et al.
Veröffentlicht: (2025)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
Bias and Fairness in Large Language Models: A Survey
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2023)
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2023)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
von: Hashmat, Abdullah, et al.
Veröffentlicht: (2025)
von: Hashmat, Abdullah, et al.
Veröffentlicht: (2025)
Impacts of Racial Bias in Historical Training Data for News AI
von: Bhargava, Rahul, et al.
Veröffentlicht: (2025)
von: Bhargava, Rahul, et al.
Veröffentlicht: (2025)
Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
von: Islam, Tunazzina
Veröffentlicht: (2026)
von: Islam, Tunazzina
Veröffentlicht: (2026)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
von: Urchs, Stefanie, et al.
Veröffentlicht: (2023)
von: Urchs, Stefanie, et al.
Veröffentlicht: (2023)
AI-AI Bias: large language models favor communications generated by large language models
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025)
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
Towards Detecting Persuasion on Social Media: From Model Development to Insights on Persuasion Strategies
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
von: Kurz, Simon, et al.
Veröffentlicht: (2024)
von: Kurz, Simon, et al.
Veröffentlicht: (2024)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
von: He, Luxi, et al.
Veröffentlicht: (2024)
von: He, Luxi, et al.
Veröffentlicht: (2024)
ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
von: Phutane, Mahika, et al.
Veröffentlicht: (2025)
von: Phutane, Mahika, et al.
Veröffentlicht: (2025)
Auto311: A Confidence-guided Automated System for Non-emergency Calls
von: Chen, Zirong, et al.
Veröffentlicht: (2023)
von: Chen, Zirong, et al.
Veröffentlicht: (2023)
Representation Bias of Adolescents in AI: A Bilingual, Bicultural Study
von: Wolfe, Robert, et al.
Veröffentlicht: (2024)
von: Wolfe, Robert, et al.
Veröffentlicht: (2024)
The Mirage of Artificial Intelligence Terms of Use Restrictions
von: Henderson, Peter, et al.
Veröffentlicht: (2024)
von: Henderson, Peter, et al.
Veröffentlicht: (2024)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
von: Wei, Boyi, et al.
Veröffentlicht: (2024)
von: Wei, Boyi, et al.
Veröffentlicht: (2024)
Representation Engineering: A Top-Down Approach to AI Transparency
von: Zou, Andy, et al.
Veröffentlicht: (2023)
von: Zou, Andy, et al.
Veröffentlicht: (2023)
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
von: Abels, Axel, et al.
Veröffentlicht: (2025)
von: Abels, Axel, et al.
Veröffentlicht: (2025)
Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models
von: Feng, Duanyu, et al.
Veröffentlicht: (2023)
von: Feng, Duanyu, et al.
Veröffentlicht: (2023)
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data
von: Sinacola, Enzo, et al.
Veröffentlicht: (2025)
von: Sinacola, Enzo, et al.
Veröffentlicht: (2025)
Multimodal Assessment of Classroom Discourse Quality: A Text-Centered Attention-Based Multi-Task Learning Approach
von: Hou, Ruikun, et al.
Veröffentlicht: (2025)
von: Hou, Ruikun, et al.
Veröffentlicht: (2025)
Which English Do LLMs Prefer? Triangulating Structural Bias Towards American English in Foundation Models
von: Nayeem, Mir Tafseer, et al.
Veröffentlicht: (2026)
von: Nayeem, Mir Tafseer, et al.
Veröffentlicht: (2026)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
Integrating Fairness and Model Pruning Through Bi-level Optimization
von: Dai, Yucong, et al.
Veröffentlicht: (2023)
von: Dai, Yucong, et al.
Veröffentlicht: (2023)
Implementing a Nordic-Baltic Federated Health Data Network: a case report
von: Chomutare, Taridzo, et al.
Veröffentlicht: (2024)
von: Chomutare, Taridzo, et al.
Veröffentlicht: (2024)
Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments
von: Feger, Marc, et al.
Veröffentlicht: (2025)
von: Feger, Marc, et al.
Veröffentlicht: (2025)
Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval
von: Wilson, Kyra, et al.
Veröffentlicht: (2024)
von: Wilson, Kyra, et al.
Veröffentlicht: (2024)
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
Dataset Scale and Societal Consistency Mediate Facial Impression Bias in Vision-Language AI
von: Wolfe, Robert, et al.
Veröffentlicht: (2024)
von: Wolfe, Robert, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
What's in a Name? Auditing Large Language Models for Race and Gender Bias
von: Salinas, Alejandro, et al.
Veröffentlicht: (2024) -
Identifying Emerging Concepts in Large Corpora
von: Ma, Sibo, et al.
Veröffentlicht: (2025) -
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
von: Xu, Xin, et al.
Veröffentlicht: (2025) -
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
von: Xu, Xin, et al.
Veröffentlicht: (2025) -
Mitigating Bias for Question Answering Models by Tracking Bias Influence
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)