Toxicity Detection for Free
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Zhanhao, Piet, Julien, Zhao, Geng, Jiao, Jiantao, Wagner, David |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generative AI Security: Challenges and Countermeasures
por: Zhu, Banghua, et al.
Publicado: (2024)
por: Zhu, Banghua, et al.
Publicado: (2024)
Mark My Words: Analyzing and Evaluating Language Model Watermarks
por: Piet, Julien, et al.
Publicado: (2023)
por: Piet, Julien, et al.
Publicado: (2023)
$k$NNProxy: Efficient Training-Free Proxy Alignment for Black-Box Zero-Shot LLM-Generated Text Detection
por: Wong, Kahim, et al.
Publicado: (2026)
por: Wong, Kahim, et al.
Publicado: (2026)
GradShield: Alignment Preserving Finetuning
por: Hu, Zhanhao, et al.
Publicado: (2026)
por: Hu, Zhanhao, et al.
Publicado: (2026)
Toward a Theory of Tokenization in LLMs
por: Rajaraman, Nived, et al.
Publicado: (2024)
por: Rajaraman, Nived, et al.
Publicado: (2024)
Efficient Prompt Caching via Embedding Similarity
por: Zhu, Hanlin, et al.
Publicado: (2024)
por: Zhu, Hanlin, et al.
Publicado: (2024)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
por: Zhao, Yibo, et al.
Publicado: (2024)
por: Zhao, Yibo, et al.
Publicado: (2024)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
por: Piet, Julien, et al.
Publicado: (2023)
por: Piet, Julien, et al.
Publicado: (2023)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
por: Zhu, Banghua, et al.
Publicado: (2024)
por: Zhu, Banghua, et al.
Publicado: (2024)
Beyond Toxic: Toxicity Detection Datasets are Not Enough for Brand Safety
por: Korotkova, Elizaveta, et al.
Publicado: (2023)
por: Korotkova, Elizaveta, et al.
Publicado: (2023)
Parser-Free Querying of Security Logs
por: Luo, Evan, et al.
Publicado: (2026)
por: Luo, Evan, et al.
Publicado: (2026)
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
por: Piet, Julien, et al.
Publicado: (2025)
por: Piet, Julien, et al.
Publicado: (2025)
Culture Matters in Toxic Language Detection in Persian
por: Bokaei, Zahra, et al.
Publicado: (2025)
por: Bokaei, Zahra, et al.
Publicado: (2025)
Translate, then Detect: Leveraging Machine Translation for Cross-Lingual Toxicity Classification
por: Bell, Samuel J., et al.
Publicado: (2025)
por: Bell, Samuel J., et al.
Publicado: (2025)
Web Agents Should Adopt the Plan-Then-Execute Paradigm
por: Piet, Julien, et al.
Publicado: (2026)
por: Piet, Julien, et al.
Publicado: (2026)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
por: Yang, Shujian, et al.
Publicado: (2025)
por: Yang, Shujian, et al.
Publicado: (2025)
Concept-Based Interpretability for Toxicity Detection
por: Garg, Samarth, et al.
Publicado: (2025)
por: Garg, Samarth, et al.
Publicado: (2025)
In-game Toxic Language Detection: Shared Task and Attention Residuals
por: Jia, Yuanzhe, et al.
Publicado: (2022)
por: Jia, Yuanzhe, et al.
Publicado: (2022)
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
por: Atil, Berk, et al.
Publicado: (2026)
por: Atil, Berk, et al.
Publicado: (2026)
Detecting Toxic Language: Ontology and BERT-based Approaches for Bulgarian Text
por: Berbatova, Melania, et al.
Publicado: (2026)
por: Berbatova, Melania, et al.
Publicado: (2026)
Finding Pareto Trade-offs in Fair and Accurate Detection of Toxic Speech
por: Gupta, Soumyajit, et al.
Publicado: (2022)
por: Gupta, Soumyajit, et al.
Publicado: (2022)
ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection
por: Li, Boyang, et al.
Publicado: (2026)
por: Li, Boyang, et al.
Publicado: (2026)
Thinking LLMs: General Instruction Following with Thought Generation
por: Wu, Tianhao, et al.
Publicado: (2024)
por: Wu, Tianhao, et al.
Publicado: (2024)
ToXCL: A Unified Framework for Toxic Speech Detection and Explanation
por: Hoang, Nhat M., et al.
Publicado: (2024)
por: Hoang, Nhat M., et al.
Publicado: (2024)
ViToSA: Audio-Based Toxic Spans Detection on Vietnamese Speech Utterances
por: Do, Huy Ba, et al.
Publicado: (2025)
por: Do, Huy Ba, et al.
Publicado: (2025)
Large Language Models for Toxic Language Detection in Low-Resource Balkan Languages
por: Muminovic, Amel, et al.
Publicado: (2025)
por: Muminovic, Amel, et al.
Publicado: (2025)
Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties
por: Faisal, Fahim, et al.
Publicado: (2024)
por: Faisal, Fahim, et al.
Publicado: (2024)
Toxicity Classification in Ukrainian
por: Dementieva, Daryna, et al.
Publicado: (2024)
por: Dementieva, Daryna, et al.
Publicado: (2024)
Benchmark on Peer Review Toxic Detection: A Challenging Task with a New Dataset
por: Luo, Man, et al.
Publicado: (2025)
por: Luo, Man, et al.
Publicado: (2025)
Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
por: Zhu, Hanlin, et al.
Publicado: (2024)
por: Zhu, Hanlin, et al.
Publicado: (2024)
Toxic Memes: A Survey of Computational Perspectives on the Detection and Explanation of Meme Toxicities
por: Pandiani, Delfina Sol Martinez, et al.
Publicado: (2024)
por: Pandiani, Delfina Sol Martinez, et al.
Publicado: (2024)
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
por: Jain, Devansh, et al.
Publicado: (2024)
por: Jain, Devansh, et al.
Publicado: (2024)
FrenchToxicityPrompts: a Large Benchmark for Evaluating and Mitigating Toxicity in French Texts
por: Brun, Caroline, et al.
Publicado: (2024)
por: Brun, Caroline, et al.
Publicado: (2024)
Context-Aware Toxicity Detection in Multiplayer Games: Integrating Domain-Adaptive Pretraining and Match Metadata
por: Schurger-Foy, Adrien, et al.
Publicado: (2025)
por: Schurger-Foy, Adrien, et al.
Publicado: (2025)
On Bias and Fairness in NLP: Investigating the Impact of Bias and Debiasing in Language Models on the Fairness of Toxicity Detection
por: Elsafoury, Fatma, et al.
Publicado: (2023)
por: Elsafoury, Fatma, et al.
Publicado: (2023)
EmbedLLM: Learning Compact Representations of Large Language Models
por: Zhuang, Richard, et al.
Publicado: (2024)
por: Zhuang, Richard, et al.
Publicado: (2024)
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
por: Hu, Yujia, et al.
Publicado: (2025)
por: Hu, Yujia, et al.
Publicado: (2025)
Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media
por: Assenmacher, Dennis, et al.
Publicado: (2025)
por: Assenmacher, Dennis, et al.
Publicado: (2025)
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
por: Duan, Zenghao, et al.
Publicado: (2025)
por: Duan, Zenghao, et al.
Publicado: (2025)
Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity
por: Uppaal, Rheeya, et al.
Publicado: (2024)
por: Uppaal, Rheeya, et al.
Publicado: (2024)
Ejemplares similares
-
Generative AI Security: Challenges and Countermeasures
por: Zhu, Banghua, et al.
Publicado: (2024) -
Mark My Words: Analyzing and Evaluating Language Model Watermarks
por: Piet, Julien, et al.
Publicado: (2023) -
$k$NNProxy: Efficient Training-Free Proxy Alignment for Black-Box Zero-Shot LLM-Generated Text Detection
por: Wong, Kahim, et al.
Publicado: (2026) -
GradShield: Alignment Preserving Finetuning
por: Hu, Zhanhao, et al.
Publicado: (2026) -
Toward a Theory of Tokenization in LLMs
por: Rajaraman, Nived, et al.
Publicado: (2024)