PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Jain, Devansh, Kumar, Priyanshu, Gehman, Samuel, Zhou, Xuhui, Hartvigsen, Thomas, Sap, Maarten |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
por: Kumar, Priyanshu, et al.
Publicado: (2025)
por: Kumar, Priyanshu, et al.
Publicado: (2025)
FrenchToxicityPrompts: a Large Benchmark for Evaluating and Mitigating Toxicity in French Texts
por: Brun, Caroline, et al.
Publicado: (2024)
por: Brun, Caroline, et al.
Publicado: (2024)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
por: Beniwal, Himanshu, et al.
Publicado: (2025)
por: Beniwal, Himanshu, et al.
Publicado: (2025)
Realistic Evaluation of Toxicity in Large Language Models
por: Luong, Tinh Son, et al.
Publicado: (2024)
por: Luong, Tinh Son, et al.
Publicado: (2024)
Efficient Detection of Toxic Prompts in Large Language Models
por: Liu, Yi, et al.
Publicado: (2024)
por: Liu, Yi, et al.
Publicado: (2024)
ToxSearch: Evolving Prompts for Toxicity Search in Large Language Models
por: Shelar, Onkar, et al.
Publicado: (2025)
por: Shelar, Onkar, et al.
Publicado: (2025)
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
por: Ghate, Kshitish, et al.
Publicado: (2025)
por: Ghate, Kshitish, et al.
Publicado: (2025)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
por: Zhou, Xuhui, et al.
Publicado: (2024)
por: Zhou, Xuhui, et al.
Publicado: (2024)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
por: Mun, Jimin, et al.
Publicado: (2026)
por: Mun, Jimin, et al.
Publicado: (2026)
RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?
por: de Wynter, Adrian, et al.
Publicado: (2024)
por: de Wynter, Adrian, et al.
Publicado: (2024)
TAXI: Evaluating Categorical Knowledge Editing for Language Models
por: Powell, Derek, et al.
Publicado: (2024)
por: Powell, Derek, et al.
Publicado: (2024)
Benchmarking LLM Guardrails in Handling Multilingual Toxicity
por: Yang, Yahan, et al.
Publicado: (2024)
por: Yang, Yahan, et al.
Publicado: (2024)
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
por: Shen, Jocelyn, et al.
Publicado: (2025)
por: Shen, Jocelyn, et al.
Publicado: (2025)
Characterising Toxicity in Generative Large Language Models
por: Zhang, Zhiyao, et al.
Publicado: (2026)
por: Zhang, Zhiyao, et al.
Publicado: (2026)
SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
por: Fan, Xianzhe, et al.
Publicado: (2025)
por: Fan, Xianzhe, et al.
Publicado: (2025)
Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English
por: Zhou, Runtao, et al.
Publicado: (2025)
por: Zhou, Runtao, et al.
Publicado: (2025)
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
por: Li, Wenkai, et al.
Publicado: (2024)
por: Li, Wenkai, et al.
Publicado: (2024)
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
por: Rao, Abhinav, et al.
Publicado: (2024)
por: Rao, Abhinav, et al.
Publicado: (2024)
Evaluating Temporal Consistency in Multi-Turn Language Models
por: Atri, Yash Kumar, et al.
Publicado: (2026)
por: Atri, Yash Kumar, et al.
Publicado: (2026)
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
por: Kou, Zhiqiang, et al.
Publicado: (2025)
por: Kou, Zhiqiang, et al.
Publicado: (2025)
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
por: Wang, Wenxuan
Publicado: (2024)
por: Wang, Wenxuan
Publicado: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
por: Mireshghallah, Niloofar, et al.
Publicado: (2023)
por: Mireshghallah, Niloofar, et al.
Publicado: (2023)
Large Language Models for Toxic Language Detection in Low-Resource Balkan Languages
por: Muminovic, Amel, et al.
Publicado: (2025)
por: Muminovic, Amel, et al.
Publicado: (2025)
Data Defenses Against Large Language Models
por: Agnew, William, et al.
Publicado: (2024)
por: Agnew, William, et al.
Publicado: (2024)
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
por: Corbo, Simone, et al.
Publicado: (2025)
por: Corbo, Simone, et al.
Publicado: (2025)
Continually Self-Improving Language Models for Bariatric Surgery Question--Answering
por: Atri, Yash Kumar, et al.
Publicado: (2025)
por: Atri, Yash Kumar, et al.
Publicado: (2025)
Leveraging Large Language Models and Topic Modeling for Toxicity Classification
por: Oskouie, Haniyeh Ehsani, et al.
Publicado: (2024)
por: Oskouie, Haniyeh Ehsani, et al.
Publicado: (2024)
Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
por: Suau, Xavier, et al.
Publicado: (2024)
por: Suau, Xavier, et al.
Publicado: (2024)
Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
por: Zhou, Kaitlyn, et al.
Publicado: (2024)
por: Zhou, Kaitlyn, et al.
Publicado: (2024)
Enhancing Multilingual Voice Toxicity Detection with Speech-Text Alignment
por: Liu, Joseph, et al.
Publicado: (2024)
por: Liu, Joseph, et al.
Publicado: (2024)
Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models
por: Lu, Hongyuan, et al.
Publicado: (2024)
por: Lu, Hongyuan, et al.
Publicado: (2024)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
por: Mendelsohn, Julia, et al.
Publicado: (2023)
por: Mendelsohn, Julia, et al.
Publicado: (2023)
MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
por: Jin, Bohan, et al.
Publicado: (2025)
por: Jin, Bohan, et al.
Publicado: (2025)
ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
por: Lee, Wonjun, et al.
Publicado: (2025)
por: Lee, Wonjun, et al.
Publicado: (2025)
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
por: Shimgekar, Soorya Ram, et al.
Publicado: (2026)
por: Shimgekar, Soorya Ram, et al.
Publicado: (2026)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
por: Beniwal, Himanshu, et al.
Publicado: (2025)
por: Beniwal, Himanshu, et al.
Publicado: (2025)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
por: Su, Zhe, et al.
Publicado: (2024)
por: Su, Zhe, et al.
Publicado: (2024)
Model Editing with Graph-Based External Memory
por: Atri, Yash Kumar, et al.
Publicado: (2025)
por: Atri, Yash Kumar, et al.
Publicado: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
por: Wang, Shuxun, et al.
Publicado: (2025)
por: Wang, Shuxun, et al.
Publicado: (2025)
Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study
por: Surana, Mokshit, et al.
Publicado: (2026)
por: Surana, Mokshit, et al.
Publicado: (2026)
Ejemplares similares
-
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
por: Kumar, Priyanshu, et al.
Publicado: (2025) -
FrenchToxicityPrompts: a Large Benchmark for Evaluating and Mitigating Toxicity in French Texts
por: Brun, Caroline, et al.
Publicado: (2024) -
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
por: Beniwal, Himanshu, et al.
Publicado: (2025) -
Realistic Evaluation of Toxicity in Large Language Models
por: Luong, Tinh Son, et al.
Publicado: (2024) -
Efficient Detection of Toxic Prompts in Large Language Models
por: Liu, Yi, et al.
Publicado: (2024)