Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
Fuente:
arXiv
Salvato in:
| Autori principali: | Ngueajio, Mikel K., Plaza-del-Arco, Flor Miriam, Chung, Yi-Ling, Rawat, Danda B., Curry, Amanda Cercas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024)
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024)
Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024)
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024)
From Chatbots to Confidants: A Cross-Cultural Study of LLM Adoption for Emotional Support
di: Amat-Lefort, Natalia, et al.
Pubblicazione: (2026)
di: Amat-Lefort, Natalia, et al.
Pubblicazione: (2026)
Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024)
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024)
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks
di: Zhao, Justin, et al.
Pubblicazione: (2024)
di: Zhao, Justin, et al.
Pubblicazione: (2024)
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation
di: Trager, Jackson, et al.
Pubblicazione: (2025)
di: Trager, Jackson, et al.
Pubblicazione: (2025)
Subjective $\textit{Isms}$? On the Danger of Conflating Hate and Offence in Abusive Language Detection
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2024)
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2024)
Impoverished Language Technology: The Lack of (Social) Class in NLP
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2024)
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2024)
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2025)
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2025)
Wisdom of Instruction-Tuned Language Model Crowds. Exploring Model Label Variation
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2023)
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2023)
The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
di: Bassignana, Elisa, et al.
Pubblicazione: (2025)
di: Bassignana, Elisa, et al.
Pubblicazione: (2025)
Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models
di: Rooein, Donya, et al.
Pubblicazione: (2025)
di: Rooein, Donya, et al.
Pubblicazione: (2025)
Classist Tools: Social Class Correlates with Performance in NLP
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2024)
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2024)
Do Large Language Models Adapt to Language Variation across Socioeconomic Status?
di: Bassignana, Elisa, et al.
Pubblicazione: (2026)
di: Bassignana, Elisa, et al.
Pubblicazione: (2026)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
di: Yuan, Shuzhou, et al.
Pubblicazione: (2025)
di: Yuan, Shuzhou, et al.
Pubblicazione: (2025)
SINAI at eRisk@CLEF 2023: Approaching Early Detection of Gambling with Natural Language Processing
di: Marmol-Romero, Alba Maria, et al.
Pubblicazione: (2025)
di: Marmol-Romero, Alba Maria, et al.
Pubblicazione: (2025)
FLANS at SemEval-2026 Task 7: RAG with Open-Sourced Smaller LLMs for Everyday Knowledge Across Diverse Languages and Cultures
di: Bogdanova, Liliia, et al.
Pubblicazione: (2026)
di: Bogdanova, Liliia, et al.
Pubblicazione: (2026)
P1SCO: Social Dimensions from a Perspectivist Lens
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2026)
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2026)
Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement
di: Abercrombie, Gavin, et al.
Pubblicazione: (2023)
di: Abercrombie, Gavin, et al.
Pubblicazione: (2023)
NLP for Counterspeech against Hate: A Survey and How-To Guide
di: Bonaldi, Helena, et al.
Pubblicazione: (2024)
di: Bonaldi, Helena, et al.
Pubblicazione: (2024)
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
di: Gajewska, Ewelina, et al.
Pubblicazione: (2025)
di: Gajewska, Ewelina, et al.
Pubblicazione: (2025)
Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering
di: Bonaldi, Helena, et al.
Pubblicazione: (2024)
di: Bonaldi, Helena, et al.
Pubblicazione: (2024)
LLMs + Persona-Plug = Personalized LLMs
di: Liu, Jiongnan, et al.
Pubblicazione: (2024)
di: Liu, Jiongnan, et al.
Pubblicazione: (2024)
The Personality Trap: How LLMs Embed Bias When Generating Human-Like Personas
di: Amidei, Jacopo, et al.
Pubblicazione: (2026)
di: Amidei, Jacopo, et al.
Pubblicazione: (2026)
Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents
di: Morocho, Erika Elizabeth Taday, et al.
Pubblicazione: (2026)
di: Morocho, Erika Elizabeth Taday, et al.
Pubblicazione: (2026)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
Dialogues of Dissent: Thematic and Rhetorical Dimensions of Hate and Counter-Hate Speech in Social Media Conversations
di: Levi, Effi, et al.
Pubblicazione: (2025)
di: Levi, Effi, et al.
Pubblicazione: (2025)
PersonaVLM: Long-Term Personalized Multimodal LLMs
di: Nie, Chang, et al.
Pubblicazione: (2026)
di: Nie, Chang, et al.
Pubblicazione: (2026)
Perceiving and Countering Hate: The Role of Identity in Online Responses
di: Ping, Kaike, et al.
Pubblicazione: (2024)
di: Ping, Kaike, et al.
Pubblicazione: (2024)
Outcome-Constrained Large Language Models for Countering Hate Speech
di: Hong, Lingzi, et al.
Pubblicazione: (2024)
di: Hong, Lingzi, et al.
Pubblicazione: (2024)
Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs
di: Tang, Wenqiu, et al.
Pubblicazione: (2026)
di: Tang, Wenqiu, et al.
Pubblicazione: (2026)
Can Thinking Models Think to Detect Hateful Memes?
di: Kmainasi, Mohamed Bayan, et al.
Pubblicazione: (2026)
di: Kmainasi, Mohamed Bayan, et al.
Pubblicazione: (2026)
SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
di: Ryan, Michael J, et al.
Pubblicazione: (2025)
di: Ryan, Michael J, et al.
Pubblicazione: (2025)
AI-Driven Human-Autonomy Teaming in Tactical Operations: Proposed Framework, Challenges, and Future Directions
di: Hagos, Desta Haileselassie, et al.
Pubblicazione: (2024)
di: Hagos, Desta Haileselassie, et al.
Pubblicazione: (2024)
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
di: Saha, Sougata, et al.
Pubblicazione: (2024)
di: Saha, Sougata, et al.
Pubblicazione: (2024)
Localizing Persona Representations in LLMs
di: Cintas, Celia, et al.
Pubblicazione: (2025)
di: Cintas, Celia, et al.
Pubblicazione: (2025)
PEACE 2.0: Grounded Explanations and Counter-Speech for Combating Hate Expressions
di: Damo, Greta, et al.
Pubblicazione: (2026)
di: Damo, Greta, et al.
Pubblicazione: (2026)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
di: Jin, Yiping, et al.
Pubblicazione: (2024)
di: Jin, Yiping, et al.
Pubblicazione: (2024)
Basque and Spanish Counter Narrative Generation: Data Creation and Evaluation
di: Bengoetxea, Jaione, et al.
Pubblicazione: (2024)
di: Bengoetxea, Jaione, et al.
Pubblicazione: (2024)
ReZG: Retrieval-Augmented Zero-Shot Counter Narrative Generation for Hate Speech
di: Jiang, Shuyu, et al.
Pubblicazione: (2023)
di: Jiang, Shuyu, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024) -
Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024) -
From Chatbots to Confidants: A Cross-Cultural Study of LLM Adoption for Emotional Support
di: Amat-Lefort, Natalia, et al.
Pubblicazione: (2026) -
Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024) -
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks
di: Zhao, Justin, et al.
Pubblicazione: (2024)