Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
Fuente:
arXiv
Saved in:
| Main Authors: | Chien, Jennifer, Danks, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
by: Shieh, Evan, et al.
Published: (2024)
by: Shieh, Evan, et al.
Published: (2024)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
Understanding and Mitigating Risks of Generative AI in Financial Services
by: Gehrmann, Sebastian, et al.
Published: (2025)
by: Gehrmann, Sebastian, et al.
Published: (2025)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
by: Sun, Shuzhou, et al.
Published: (2025)
by: Sun, Shuzhou, et al.
Published: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
by: Arghal, Raghu, et al.
Published: (2026)
by: Arghal, Raghu, et al.
Published: (2026)
EigenBench: A Comparative Behavioral Measure of Value Alignment
by: Chang, Jonathn, et al.
Published: (2025)
by: Chang, Jonathn, et al.
Published: (2025)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
AWARE, Beyond Sentence Boundaries: A Contextual Transformer Framework for Identifying Cultural Capital in STEM Narratives
by: Khan, Khalid Mehtab, et al.
Published: (2025)
by: Khan, Khalid Mehtab, et al.
Published: (2025)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
by: Hu, Zhanghao, et al.
Published: (2025)
by: Hu, Zhanghao, et al.
Published: (2025)
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Computational Measurement of Political Positions: A Review of Text-Based Ideal Point Estimation Algorithms
by: Parschan, Patrick, et al.
Published: (2025)
by: Parschan, Patrick, et al.
Published: (2025)
Language Representation Favored Zero-Shot Cross-Domain Cognitive Diagnosis
by: Liu, Shuo, et al.
Published: (2025)
by: Liu, Shuo, et al.
Published: (2025)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
by: Arakelyan, Erik, et al.
Published: (2024)
by: Arakelyan, Erik, et al.
Published: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
CEGI: Measuring the trade-off between efficiency and carbon emissions for SLMs and VLMs
by: Kumar, Abhas, et al.
Published: (2024)
by: Kumar, Abhas, et al.
Published: (2024)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
Do Large Language Models Walk Their Talk? Measuring the Gap Between Implicit Associations, Self-Report, and Behavioral Altruism
by: Andric, Sandro
Published: (2025)
by: Andric, Sandro
Published: (2025)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
by: Chehbouni, Khaoula, et al.
Published: (2024)
by: Chehbouni, Khaoula, et al.
Published: (2024)
Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
by: Cunningham, Eoghan, et al.
Published: (2025)
by: Cunningham, Eoghan, et al.
Published: (2025)
LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
by: Bouchard, Dylan, et al.
Published: (2025)
by: Bouchard, Dylan, et al.
Published: (2025)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025)
by: Dobre, David, et al.
Published: (2025)
Large Language Models Predict Functional Outcomes after Acute Ischemic Stroke
by: Kapoor, Anjali K., et al.
Published: (2026)
by: Kapoor, Anjali K., et al.
Published: (2026)
Representation Bias of Adolescents in AI: A Bilingual, Bicultural Study
by: Wolfe, Robert, et al.
Published: (2024)
by: Wolfe, Robert, et al.
Published: (2024)
Large Language Models are Geographically Biased
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
by: Berezin, Sergei, et al.
Published: (2025)
by: Berezin, Sergei, et al.
Published: (2025)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
by: Abels, Axel, et al.
Published: (2025)
by: Abels, Axel, et al.
Published: (2025)
Harm Amplification in Text-to-Image Models
by: Hao, Susan, et al.
Published: (2024)
by: Hao, Susan, et al.
Published: (2024)
Towards Best Practices for Open Datasets for LLM Training
by: Baack, Stefan, et al.
Published: (2025)
by: Baack, Stefan, et al.
Published: (2025)
Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
Beyond Via: Analysis and Estimation of the Impact of Large Language Models in Academic Papers
by: Geng, Mingmeng, et al.
Published: (2026)
by: Geng, Mingmeng, et al.
Published: (2026)
How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
by: Schroeder, Daniel Thilo, et al.
Published: (2025)
by: Schroeder, Daniel Thilo, et al.
Published: (2025)
A Multi-LLM Debiasing Framework
by: Owens, Deonna M., et al.
Published: (2024)
by: Owens, Deonna M., et al.
Published: (2024)
A Taxonomy of Stereotype Content in Large Language Models
by: Nicolas, Gandalf, et al.
Published: (2024)
by: Nicolas, Gandalf, et al.
Published: (2024)
Linear Representations of Political Perspective Emerge in Large Language Models
by: Kim, Junsol, et al.
Published: (2025)
by: Kim, Junsol, et al.
Published: (2025)
Similar Items
-
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026) -
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
by: Shieh, Evan, et al.
Published: (2024) -
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025) -
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025) -
Understanding and Mitigating Risks of Generative AI in Financial Services
by: Gehrmann, Sebastian, et al.
Published: (2025)