PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Menke, Maluna, Hagendorff, Thilo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025)
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025)
On the Inevitability of Left-Leaning Political Bias in Aligned Language Models
von: Hagendorff, Thilo
Veröffentlicht: (2025)
von: Hagendorff, Thilo
Veröffentlicht: (2025)
Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2025)
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2025)
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
von: Knecht, Amelie, et al.
Veröffentlicht: (2026)
von: Knecht, Amelie, et al.
Veröffentlicht: (2026)
Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
von: Hagendorff, Thilo
Veröffentlicht: (2024)
von: Hagendorff, Thilo
Veröffentlicht: (2024)
When Image Generation Goes Wrong: A Safety Analysis of Stable Diffusion Models
von: Schneider, Matthias, et al.
Veröffentlicht: (2024)
von: Schneider, Matthias, et al.
Veröffentlicht: (2024)
Deception Abilities Emerged in Large Language Models
von: Hagendorff, Thilo
Veröffentlicht: (2023)
von: Hagendorff, Thilo
Veröffentlicht: (2023)
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
Fairness Hacking: The Malicious Practice of Shrouding Unfairness in Algorithms
von: Meding, Kristof, et al.
Veröffentlicht: (2023)
von: Meding, Kristof, et al.
Veröffentlicht: (2023)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
QueerBench: Quantifying Discrimination in Language Models Toward Queer Identities
von: Sosto, Mae, et al.
Veröffentlicht: (2024)
von: Sosto, Mae, et al.
Veröffentlicht: (2024)
Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2026)
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2026)
A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2024)
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2024)
Large Reasoning Models Are Autonomous Jailbreak Agents
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
von: Fleisig, Eve, et al.
Veröffentlicht: (2024)
von: Fleisig, Eve, et al.
Veröffentlicht: (2024)
Fair Play in the Newsroom: Actor-Based Filtering Gender Discrimination in Text Corpora
von: Urchs, Stefanie, et al.
Veröffentlicht: (2025)
von: Urchs, Stefanie, et al.
Veröffentlicht: (2025)
Focal Inferential Infusion Coupled with Tractable Density Discrimination for Implicit Hate Detection
von: Masud, Sarah, et al.
Veröffentlicht: (2023)
von: Masud, Sarah, et al.
Veröffentlicht: (2023)
Examining Identity Drift in Conversations of LLM Agents
von: Choi, Junhyuk, et al.
Veröffentlicht: (2024)
von: Choi, Junhyuk, et al.
Veröffentlicht: (2024)
Clinical Note Bloat Reduction for Efficient LLM Use
von: Cahoon, Jordan L., et al.
Veröffentlicht: (2026)
von: Cahoon, Jordan L., et al.
Veröffentlicht: (2026)
Whose Journey Matters? Investigating Identity Biases in Large Language Models (LLMs) for Travel Planning Assistance
von: Ren, Ruiping, et al.
Veröffentlicht: (2024)
von: Ren, Ruiping, et al.
Veröffentlicht: (2024)
Generative Language Models Exhibit Social Identity Biases
von: Hu, Tiancheng, et al.
Veröffentlicht: (2023)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2023)
The Statistical Signature of LLMs
von: Hadad, Ortal, et al.
Veröffentlicht: (2026)
von: Hadad, Ortal, et al.
Veröffentlicht: (2026)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
von: Schmucker, Robin, et al.
Veröffentlicht: (2025)
von: Schmucker, Robin, et al.
Veröffentlicht: (2025)
Urban Mobility Assessment Using LLMs
von: Bhandari, Prabin, et al.
Veröffentlicht: (2024)
von: Bhandari, Prabin, et al.
Veröffentlicht: (2024)
Challenger at MultiPRIDE: Is It Hate Speech or Reclaimed?
von: Tekanlou, Hadi Bayrami Asl, et al.
Veröffentlicht: (2026)
von: Tekanlou, Hadi Bayrami Asl, et al.
Veröffentlicht: (2026)
The Thin Line Between Comprehension and Persuasion in LLMs
von: de Wynter, Adrian, et al.
Veröffentlicht: (2025)
von: de Wynter, Adrian, et al.
Veröffentlicht: (2025)
LLMs Provide Unstable Answers to Legal Questions
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
Academically intelligent LLMs are not necessarily socially intelligent
von: Xu, Ruoxi, et al.
Veröffentlicht: (2024)
von: Xu, Ruoxi, et al.
Veröffentlicht: (2024)
Unsupervised Concept Vector Extraction for Bias Control in LLMs
von: Cyberey, Hannah, et al.
Veröffentlicht: (2025)
von: Cyberey, Hannah, et al.
Veröffentlicht: (2025)
Surface Reading LLMs: Synthetic Text and its Styles
von: Bajohr, Hannes
Veröffentlicht: (2025)
von: Bajohr, Hannes
Veröffentlicht: (2025)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
von: Potter, Yujin, et al.
Veröffentlicht: (2024)
von: Potter, Yujin, et al.
Veröffentlicht: (2024)
Hate Personified: Investigating the role of LLMs in content moderation
von: Masud, Sarah, et al.
Veröffentlicht: (2024)
von: Masud, Sarah, et al.
Veröffentlicht: (2024)
Prompt Refinement or Fine-tuning? Best Practices for using LLMs in Computational Social Science Tasks
von: Møller, Anders Giovanni, et al.
Veröffentlicht: (2024)
von: Møller, Anders Giovanni, et al.
Veröffentlicht: (2024)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
von: Hartmann, David, et al.
Veröffentlicht: (2026)
von: Hartmann, David, et al.
Veröffentlicht: (2026)
Exploring LLMs for Predicting Tutor Strategy and Student Outcomes in Dialogues
von: Ikram, Fareya, et al.
Veröffentlicht: (2025)
von: Ikram, Fareya, et al.
Veröffentlicht: (2025)
Evaluating the Simulation of Human Personality-Driven Susceptibility to Misinformation with LLMs
von: Pratelli, Manuel, et al.
Veröffentlicht: (2025)
von: Pratelli, Manuel, et al.
Veröffentlicht: (2025)
Measuring Fine-Grained Negotiation Tactics of Humans and LLMs in Diplomacy
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
Trust, Safety, and Accuracy: Assessing LLMs for Routine Maternity Advice
von: Divya, V Sai, et al.
Veröffentlicht: (2026)
von: Divya, V Sai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025) -
On the Inevitability of Left-Leaning Political Bias in Aligned Language Models
von: Hagendorff, Thilo
Veröffentlicht: (2025) -
Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2025) -
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
von: Knecht, Amelie, et al.
Veröffentlicht: (2026) -
Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
von: Hagendorff, Thilo
Veröffentlicht: (2024)