On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
Fuente:
arXiv
Saved in:
| Main Authors: | Ghorbanpour, Faeze, Fraser, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
by: Niazi, Ruhallah, et al.
Published: (2026)
by: Niazi, Ruhallah, et al.
Published: (2026)
Are BabyLMs Second Language Learners?
by: Edman, Lukas, et al.
Published: (2024)
by: Edman, Lukas, et al.
Published: (2024)
Hate Personified: Investigating the role of LLMs in content moderation
by: Masud, Sarah, et al.
Published: (2024)
by: Masud, Sarah, et al.
Published: (2024)
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
by: Arnaiz-Rodriguez, Adrian, et al.
Published: (2025)
by: Arnaiz-Rodriguez, Adrian, et al.
Published: (2025)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
by: Cyberey, Hannah, et al.
Published: (2024)
by: Cyberey, Hannah, et al.
Published: (2024)
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation
by: Neumann, Terrence, et al.
Published: (2024)
by: Neumann, Terrence, et al.
Published: (2024)
HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
by: Wang, Kaixuan, et al.
Published: (2025)
by: Wang, Kaixuan, et al.
Published: (2025)
Prompt Refinement or Fine-tuning? Best Practices for using LLMs in Computational Social Science Tasks
by: Møller, Anders Giovanni, et al.
Published: (2024)
by: Møller, Anders Giovanni, et al.
Published: (2024)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
by: Mohamadi, Alireza, et al.
Published: (2025)
by: Mohamadi, Alireza, et al.
Published: (2025)
QueerGen: How LLMs Reflect Societal Norms on Gender and Sexuality in Sentence Completion Tasks
by: Sosto, Mae, et al.
Published: (2026)
by: Sosto, Mae, et al.
Published: (2026)
Careless Whisper: Speech-to-Text Hallucination Harms
by: Koenecke, Allison, et al.
Published: (2024)
by: Koenecke, Allison, et al.
Published: (2024)
Differentiating Emigration from Return Migration of Scholars Using Name-Based Nationality Detection Models
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025)
by: Corvi, Emily, et al.
Published: (2025)
LLM-based Semantic Augmentation for Harmful Content Detection
by: Meguellati, Elyas, et al.
Published: (2025)
by: Meguellati, Elyas, et al.
Published: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
A Capabilities Approach to Studying Bias and Harm in Language Technologies
by: Nigatu, Hellina Hailu, et al.
Published: (2024)
by: Nigatu, Hellina Hailu, et al.
Published: (2024)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
by: Dorn, Rebecca, et al.
Published: (2024)
by: Dorn, Rebecca, et al.
Published: (2024)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
by: Chehbouni, Khaoula, et al.
Published: (2024)
by: Chehbouni, Khaoula, et al.
Published: (2024)
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers
by: Chakrabarty, Tuhin, et al.
Published: (2023)
by: Chakrabarty, Tuhin, et al.
Published: (2023)
From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2025)
by: Harvey, Emma, et al.
Published: (2025)
Exploring LLMs for Predicting Tutor Strategy and Student Outcomes in Dialogues
by: Ikram, Fareya, et al.
Published: (2025)
by: Ikram, Fareya, et al.
Published: (2025)
Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods
by: Fraser, Kathleen C., et al.
Published: (2024)
by: Fraser, Kathleen C., et al.
Published: (2024)
What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text
by: Kanepajs, Arturs, et al.
Published: (2025)
by: Kanepajs, Arturs, et al.
Published: (2025)
Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
by: Gajewska, Ewelina, et al.
Published: (2026)
by: Gajewska, Ewelina, et al.
Published: (2026)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
by: Sen, Indira, et al.
Published: (2023)
by: Sen, Indira, et al.
Published: (2023)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024)
by: Jin, Yiping, et al.
Published: (2024)
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Fine-tuning with Hierarchical Prompting for Robust Propaganda Classification Across Annotation Schemas
by: Stähelin, Lukas, et al.
Published: (2026)
by: Stähelin, Lukas, et al.
Published: (2026)
Topic Modeling with Fine-tuning LLMs and Bag of Sentences
by: Schneider, Johannes
Published: (2024)
by: Schneider, Johannes
Published: (2024)
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
by: Zhang, Yazhou, et al.
Published: (2025)
by: Zhang, Yazhou, et al.
Published: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Similar Items
-
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
by: Ghorbanpour, Faeze, et al.
Published: (2025) -
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
by: Ghorbanpour, Faeze, et al.
Published: (2025) -
PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
by: Niazi, Ruhallah, et al.
Published: (2026) -
Are BabyLMs Second Language Learners?
by: Edman, Lukas, et al.
Published: (2024) -
Hate Personified: Investigating the role of LLMs in content moderation
by: Masud, Sarah, et al.
Published: (2024)