LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
Fuente:
arXiv
Saved in:
| Main Authors: | Foo, Jessica, Khoo, Shaun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
by: Tan, Leanne, et al.
Published: (2025)
by: Tan, Leanne, et al.
Published: (2025)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
by: Wróbel, Krzysztof, et al.
Published: (2026)
by: Wróbel, Krzysztof, et al.
Published: (2026)
Safe at the Margins: A General Approach to Safety Alignment in Low-Resource English Languages -- A Singlish Case Study
by: Lim, Isaac, et al.
Published: (2025)
by: Lim, Isaac, et al.
Published: (2025)
MinorBench: A hand-built benchmark for content-based risks for children
by: Khoo, Shaun, et al.
Published: (2025)
by: Khoo, Shaun, et al.
Published: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
by: Elesedy, Hayder, et al.
Published: (2024)
by: Elesedy, Hayder, et al.
Published: (2024)
Know Or Not: a library for evaluating out-of-knowledge base robustness
by: Foo, Jessica, et al.
Published: (2025)
by: Foo, Jessica, et al.
Published: (2025)
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
by: Khoo, Shaun, et al.
Published: (2025)
by: Khoo, Shaun, et al.
Published: (2025)
Contextual Evaluation of Large Language Models for Classifying Tropical and Infectious Diseases
by: Asiedu, Mercy, et al.
Published: (2024)
by: Asiedu, Mercy, et al.
Published: (2024)
Context-Aware Content Moderation for German Newspaper Comments
by: Krejca, Felix, et al.
Published: (2025)
by: Krejca, Felix, et al.
Published: (2025)
Building Efficient Universal Classifiers with Natural Language Inference
by: Laurer, Moritz, et al.
Published: (2023)
by: Laurer, Moritz, et al.
Published: (2023)
WebGuard: Building a Generalizable Guardrail for Web Agents
by: Zheng, Boyuan, et al.
Published: (2025)
by: Zheng, Boyuan, et al.
Published: (2025)
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
by: Achara, Akshit, et al.
Published: (2025)
by: Achara, Akshit, et al.
Published: (2025)
IPS: In-Prompt Process Supervision for Short Video Content Moderation
by: Liu, Mingchao, et al.
Published: (2024)
by: Liu, Mingchao, et al.
Published: (2024)
Class-RAG: Real-Time Content Moderation with Retrieval Augmented Generation
by: Chen, Jianfa, et al.
Published: (2024)
by: Chen, Jianfa, et al.
Published: (2024)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
Gradient Co-occurrence Analysis for Detecting Unsafe Prompts in Large Language Models
by: Yang, Jingyuan, et al.
Published: (2025)
by: Yang, Jingyuan, et al.
Published: (2025)
WangchanLion and WangchanX MRC Eval
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2024)
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2024)
Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation
by: Zhang, Shutong, et al.
Published: (2026)
by: Zhang, Shutong, et al.
Published: (2026)
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
by: She, Yining, et al.
Published: (2025)
by: She, Yining, et al.
Published: (2025)
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
by: Zhuang, Jun, et al.
Published: (2025)
by: Zhuang, Jun, et al.
Published: (2025)
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators
by: Cao, Yang Trista, et al.
Published: (2023)
by: Cao, Yang Trista, et al.
Published: (2023)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
by: Yuan, Youliang, et al.
Published: (2024)
by: Yuan, Youliang, et al.
Published: (2024)
Tackling the Inherent Difficulty of Noise Filtering in RAG
by: Liu, Jingyu, et al.
Published: (2026)
by: Liu, Jingyu, et al.
Published: (2026)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
Metamorphic Testing for Audio Content Moderation Software
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Neural Contextual Reinforcement Framework for Logical Structure Language Generation
by: Irvin, Marcus, et al.
Published: (2025)
by: Irvin, Marcus, et al.
Published: (2025)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
by: Dinuta, Eduard Stefan, et al.
Published: (2025)
by: Dinuta, Eduard Stefan, et al.
Published: (2025)
X-Guard: Multilingual Guard Agent for Content Moderation
by: Upadhayay, Bibek, et al.
Published: (2025)
by: Upadhayay, Bibek, et al.
Published: (2025)
From Perceptions To Evidence: Detecting AI-Generated Content In Turkish News Media With A Fine-Tuned Bert Classifier
by: Ozdemir, Ozancan
Published: (2026)
by: Ozdemir, Ozancan
Published: (2026)
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content
by: Furniturewala, Shaz, et al.
Published: (2025)
by: Furniturewala, Shaz, et al.
Published: (2025)
Guarding the Meaning: Self-Supervised Training for Semantic Robustness in Guard Models
by: Pinneri, Cristina, et al.
Published: (2025)
by: Pinneri, Cristina, et al.
Published: (2025)
Understanding and Tackling Label Errors in Individual-Level Nature Language Understanding
by: Xiao, Yunpeng, et al.
Published: (2025)
by: Xiao, Yunpeng, et al.
Published: (2025)
Pseudocode-Injection Magic: Enabling LLMs to Tackle Graph Computational Tasks
by: Gong, Chang, et al.
Published: (2025)
by: Gong, Chang, et al.
Published: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026)
by: Tu, Xinming, et al.
Published: (2026)
AI Content Moderation in Therapy Conversations
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
by: Giarrusso, Francesco, et al.
Published: (2025)
by: Giarrusso, Francesco, et al.
Published: (2025)
The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
by: Li, Zichao, et al.
Published: (2025)
by: Li, Zichao, et al.
Published: (2025)
FactGuard: Event-Centric and Commonsense-Guided Fake News Detection
by: He, Jing, et al.
Published: (2025)
by: He, Jing, et al.
Published: (2025)
Similar Items
-
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
by: Tan, Leanne, et al.
Published: (2025) -
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
by: Wróbel, Krzysztof, et al.
Published: (2026) -
Safe at the Margins: A General Approach to Safety Alignment in Low-Resource English Languages -- A Singlish Case Study
by: Lim, Isaac, et al.
Published: (2025) -
MinorBench: A hand-built benchmark for content-based risks for children
by: Khoo, Shaun, et al.
Published: (2025) -
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)