HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Kaixuan, Diao, Chenxin, Jacques, Jason T., Guo, Zhongliang, Zhao, Shuai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Critical Challenges in Content Moderation for People Who Use Drugs (PWUD): Insights into Online Harm Reduction Practices from Moderators
por: Wang, Kaixuan, et al.
Publicado: (2025)
por: Wang, Kaixuan, et al.
Publicado: (2025)
Positioning AI Tools to Support Online Harm Reduction Practice: Applications and Design Directions
por: Wang, Kaixuan, et al.
Publicado: (2025)
por: Wang, Kaixuan, et al.
Publicado: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
por: Li, Jing-Jing, et al.
Publicado: (2026)
por: Li, Jing-Jing, et al.
Publicado: (2026)
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
por: Ghorbanpour, Faeze, et al.
Publicado: (2025)
por: Ghorbanpour, Faeze, et al.
Publicado: (2025)
Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
por: Gajewska, Ewelina, et al.
Publicado: (2026)
por: Gajewska, Ewelina, et al.
Publicado: (2026)
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
por: Arnaiz-Rodriguez, Adrian, et al.
Publicado: (2025)
por: Arnaiz-Rodriguez, Adrian, et al.
Publicado: (2025)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
por: Cyberey, Hannah, et al.
Publicado: (2024)
por: Cyberey, Hannah, et al.
Publicado: (2024)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
por: Sen, Indira, et al.
Publicado: (2023)
por: Sen, Indira, et al.
Publicado: (2023)
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation
por: Neumann, Terrence, et al.
Publicado: (2024)
por: Neumann, Terrence, et al.
Publicado: (2024)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
por: Mohamadi, Alireza, et al.
Publicado: (2025)
por: Mohamadi, Alireza, et al.
Publicado: (2025)
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
por: Menke, Maluna, et al.
Publicado: (2025)
por: Menke, Maluna, et al.
Publicado: (2025)
LLM-Generated Feedback Supports Learning If Learners Choose to Use It
por: Thomas, Danielle R., et al.
Publicado: (2025)
por: Thomas, Danielle R., et al.
Publicado: (2025)
Careless Whisper: Speech-to-Text Hallucination Harms
por: Koenecke, Allison, et al.
Publicado: (2024)
por: Koenecke, Allison, et al.
Publicado: (2024)
Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use
por: Chandra, Mohit, et al.
Publicado: (2024)
por: Chandra, Mohit, et al.
Publicado: (2024)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
por: Cheng, Myra, et al.
Publicado: (2026)
por: Cheng, Myra, et al.
Publicado: (2026)
Taxonomizing Representational Harms using Speech Act Theory
por: Corvi, Emily, et al.
Publicado: (2025)
por: Corvi, Emily, et al.
Publicado: (2025)
LLM-based Semantic Augmentation for Harmful Content Detection
por: Meguellati, Elyas, et al.
Publicado: (2025)
por: Meguellati, Elyas, et al.
Publicado: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
por: Choi, Sooyung, et al.
Publicado: (2025)
por: Choi, Sooyung, et al.
Publicado: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
por: Chen, Yen-Shan, et al.
Publicado: (2026)
por: Chen, Yen-Shan, et al.
Publicado: (2026)
Who is Undercover? Guiding LLMs to Explore Multi-Perspective Team Tactic in the Game
por: Dong, Ruiqi, et al.
Publicado: (2024)
por: Dong, Ruiqi, et al.
Publicado: (2024)
A Capabilities Approach to Studying Bias and Harm in Language Technologies
por: Nigatu, Hellina Hailu, et al.
Publicado: (2024)
por: Nigatu, Hellina Hailu, et al.
Publicado: (2024)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
por: Zhou, Yuhang, et al.
Publicado: (2025)
por: Zhou, Yuhang, et al.
Publicado: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
por: Chan, Yik Siu, et al.
Publicado: (2025)
por: Chan, Yik Siu, et al.
Publicado: (2025)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
por: Xiao, Yang, et al.
Publicado: (2023)
por: Xiao, Yang, et al.
Publicado: (2023)
Clinical Note Bloat Reduction for Efficient LLM Use
por: Cahoon, Jordan L., et al.
Publicado: (2026)
por: Cahoon, Jordan L., et al.
Publicado: (2026)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
por: Li, Chance Jiajie, et al.
Publicado: (2025)
por: Li, Chance Jiajie, et al.
Publicado: (2025)
SMILE: Single-turn to Multi-turn Inclusive Language Expansion via ChatGPT for Mental Health Support
por: Qiu, Huachuan, et al.
Publicado: (2023)
por: Qiu, Huachuan, et al.
Publicado: (2023)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
por: Dorn, Rebecca, et al.
Publicado: (2024)
por: Dorn, Rebecca, et al.
Publicado: (2024)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
por: Zhang, Chi, et al.
Publicado: (2025)
por: Zhang, Chi, et al.
Publicado: (2025)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
por: Chehbouni, Khaoula, et al.
Publicado: (2024)
por: Chehbouni, Khaoula, et al.
Publicado: (2024)
Benchmarking LLMs for Political Science: A United Nations Perspective
por: Liang, Yueqing, et al.
Publicado: (2025)
por: Liang, Yueqing, et al.
Publicado: (2025)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
por: Allaham, Mowafak, et al.
Publicado: (2024)
por: Allaham, Mowafak, et al.
Publicado: (2024)
From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
por: Li, Yang, et al.
Publicado: (2025)
por: Li, Yang, et al.
Publicado: (2025)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
por: Harvey, Emma, et al.
Publicado: (2025)
por: Harvey, Emma, et al.
Publicado: (2025)
What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text
por: Kanepajs, Arturs, et al.
Publicado: (2025)
por: Kanepajs, Arturs, et al.
Publicado: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
por: Majumdar, Ayan, et al.
Publicado: (2025)
por: Majumdar, Ayan, et al.
Publicado: (2025)
Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making
por: Drinkall, Toby
Publicado: (2025)
por: Drinkall, Toby
Publicado: (2025)
Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues
por: Duan, Zhangqi, et al.
Publicado: (2026)
por: Duan, Zhangqi, et al.
Publicado: (2026)
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
por: Maurer, Maximilian, et al.
Publicado: (2026)
por: Maurer, Maximilian, et al.
Publicado: (2026)
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models
por: Siddique, Zara, et al.
Publicado: (2024)
por: Siddique, Zara, et al.
Publicado: (2024)
Ejemplares similares
-
Critical Challenges in Content Moderation for People Who Use Drugs (PWUD): Insights into Online Harm Reduction Practices from Moderators
por: Wang, Kaixuan, et al.
Publicado: (2025) -
Positioning AI Tools to Support Online Harm Reduction Practice: Applications and Design Directions
por: Wang, Kaixuan, et al.
Publicado: (2025) -
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
por: Li, Jing-Jing, et al.
Publicado: (2026) -
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
por: Ghorbanpour, Faeze, et al.
Publicado: (2025) -
Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
por: Gajewska, Ewelina, et al.
Publicado: (2026)