Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Koh, Hyukhun, Kim, Dohyung, Lee, Minwoo, Jung, Kyomin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Conditional [MASK] Discrete Diffusion Language Model
von: Koh, Hyukhun, et al.
Veröffentlicht: (2024)
von: Koh, Hyukhun, et al.
Veröffentlicht: (2024)
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
Public Data Assisted Differentially Private In-Context Learning
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
Fine-grained Gender Control in Machine Translation with Large Language Models
von: Lee, Minwoo, et al.
Veröffentlicht: (2024)
von: Lee, Minwoo, et al.
Veröffentlicht: (2024)
Confidence-Guided Stepwise Model Routing for Cost-Efficient Reasoning
von: Lee, Sangmook, et al.
Veröffentlicht: (2025)
von: Lee, Sangmook, et al.
Veröffentlicht: (2025)
Program Synthesis via Test-Time Transduction
von: Lee, Kang-il, et al.
Veröffentlicht: (2025)
von: Lee, Kang-il, et al.
Veröffentlicht: (2025)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
When Wording Steers the Evaluation: Framing Bias in LLM judges
von: Hwang, Yerin, et al.
Veröffentlicht: (2026)
von: Hwang, Yerin, et al.
Veröffentlicht: (2026)
Casual as an Anchor: Resolving Supervision Misalignment in Formality Transfer Dataset
von: Yu, Hyojeong, et al.
Veröffentlicht: (2026)
von: Yu, Hyojeong, et al.
Veröffentlicht: (2026)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
von: Kim, Jeonghye, et al.
Veröffentlicht: (2025)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2025)
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
von: Duan, Zenghao, et al.
Veröffentlicht: (2025)
von: Duan, Zenghao, et al.
Veröffentlicht: (2025)
FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge
von: Yang, Nakyeong, et al.
Veröffentlicht: (2025)
von: Yang, Nakyeong, et al.
Veröffentlicht: (2025)
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
von: Corbo, Simone, et al.
Veröffentlicht: (2025)
von: Corbo, Simone, et al.
Veröffentlicht: (2025)
LLMs can be easily Confused by Instructional Distractions
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
Beyond Toxic: Toxicity Detection Datasets are Not Enough for Brand Safety
von: Korotkova, Elizaveta, et al.
Veröffentlicht: (2023)
von: Korotkova, Elizaveta, et al.
Veröffentlicht: (2023)
Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts
von: Rizwan, Naquee, et al.
Veröffentlicht: (2025)
von: Rizwan, Naquee, et al.
Veröffentlicht: (2025)
Generating Diverse Hypotheses for Inductive Reasoning
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
von: Zhao, Yibo, et al.
Veröffentlicht: (2024)
von: Zhao, Yibo, et al.
Veröffentlicht: (2024)
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
von: Ma, Xuchen, et al.
Veröffentlicht: (2025)
von: Ma, Xuchen, et al.
Veröffentlicht: (2025)
Building Resource-Constrained Language Agents: A Korean Case Study on Chemical Toxicity Information
von: Cho, Hojun, et al.
Veröffentlicht: (2025)
von: Cho, Hojun, et al.
Veröffentlicht: (2025)
LifeTox: Unveiling Implicit Toxicity in Life Advice
von: Kim, Minbeom, et al.
Veröffentlicht: (2023)
von: Kim, Minbeom, et al.
Veröffentlicht: (2023)
IndoToxic2024: A Demographically-Enriched Dataset of Hate Speech and Toxicity Types for Indonesian Language
von: Susanto, Lucky, et al.
Veröffentlicht: (2024)
von: Susanto, Lucky, et al.
Veröffentlicht: (2024)
Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs
von: Hwang, Kyomin, et al.
Veröffentlicht: (2025)
von: Hwang, Kyomin, et al.
Veröffentlicht: (2025)
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs
von: Hwang, Yerin, et al.
Veröffentlicht: (2024)
von: Hwang, Yerin, et al.
Veröffentlicht: (2024)
LongStory: Coherent, Complete and Length Controlled Long story Generation
von: Park, Kyeongman, et al.
Veröffentlicht: (2023)
von: Park, Kyeongman, et al.
Veröffentlicht: (2023)
Avoidance Decoding for Diverse Multi-Branch Story Generation
von: Park, Kyeongman, et al.
Veröffentlicht: (2025)
von: Park, Kyeongman, et al.
Veröffentlicht: (2025)
MVMR: A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple Distractors
von: Yang, Nakyeong, et al.
Veröffentlicht: (2023)
von: Yang, Nakyeong, et al.
Veröffentlicht: (2023)
Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic Language
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
Concept-Based Interpretability for Toxicity Detection
von: Garg, Samarth, et al.
Veröffentlicht: (2025)
von: Garg, Samarth, et al.
Veröffentlicht: (2025)
ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations
von: Kang, Hankun, et al.
Veröffentlicht: (2026)
von: Kang, Hankun, et al.
Veröffentlicht: (2026)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
von: Lee, Andrew, et al.
Veröffentlicht: (2024)
von: Lee, Andrew, et al.
Veröffentlicht: (2024)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
von: Hui, Zheng, et al.
Veröffentlicht: (2024)
von: Hui, Zheng, et al.
Veröffentlicht: (2024)
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions
von: Xu, Jingxin, et al.
Veröffentlicht: (2025)
von: Xu, Jingxin, et al.
Veröffentlicht: (2025)
Unplug and Play Language Models: Decomposing Experts in Language Models at Inference Time
von: Yang, Nakyeong, et al.
Veröffentlicht: (2024)
von: Yang, Nakyeong, et al.
Veröffentlicht: (2024)
Realistic Evaluation of Toxicity in Large Language Models
von: Luong, Tinh Son, et al.
Veröffentlicht: (2024)
von: Luong, Tinh Son, et al.
Veröffentlicht: (2024)
Characterising Toxicity in Generative Large Language Models
von: Zhang, Zhiyao, et al.
Veröffentlicht: (2026)
von: Zhang, Zhiyao, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
FMI@SU ToxHabits: Evaluating LLMs Performance on Toxic Habit Extraction in Spanish Clinical Texts
von: Vassileva, Sylvia, et al.
Veröffentlicht: (2026)
von: Vassileva, Sylvia, et al.
Veröffentlicht: (2026)
Toxic Memes: A Survey of Computational Perspectives on the Detection and Explanation of Meme Toxicities
von: Pandiani, Delfina Sol Martinez, et al.
Veröffentlicht: (2024)
von: Pandiani, Delfina Sol Martinez, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Conditional [MASK] Discrete Diffusion Language Model
von: Koh, Hyukhun, et al.
Veröffentlicht: (2024) -
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
von: Joo, Seongho, et al.
Veröffentlicht: (2025) -
Public Data Assisted Differentially Private In-Context Learning
von: Joo, Seongho, et al.
Veröffentlicht: (2025) -
Fine-grained Gender Control in Machine Translation with Large Language Models
von: Lee, Minwoo, et al.
Veröffentlicht: (2024) -
Confidence-Guided Stepwise Model Routing for Cost-Efficient Reasoning
von: Lee, Sangmook, et al.
Veröffentlicht: (2025)