LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Leanne, Chua, Gabriel, Ge, Ziyu, Lee, Roy Ka-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
by: Chua, Gabriel, et al.
Published: (2025)
by: Chua, Gabriel, et al.
Published: (2025)
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024)
by: Foo, Jessica, et al.
Published: (2024)
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
by: Ge, Ziyu, et al.
Published: (2025)
by: Ge, Ziyu, et al.
Published: (2025)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
by: Liu, Hongfu, et al.
Published: (2024)
by: Liu, Hongfu, et al.
Published: (2024)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
by: Elesedy, Hayder, et al.
Published: (2024)
by: Elesedy, Hayder, et al.
Published: (2024)
CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
by: Joshi, Raviraj, et al.
Published: (2025)
by: Joshi, Raviraj, et al.
Published: (2025)
DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails
by: Deng, Yihe, et al.
Published: (2025)
by: Deng, Yihe, et al.
Published: (2025)
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Contrastive Token-level Explanations for Graph-based Rumour Detection
by: Chin, Daniel Wai Kit, et al.
Published: (2025)
by: Chin, Daniel Wai Kit, et al.
Published: (2025)
ShieldGemma: Generative AI Content Moderation Based on Gemma
by: Zeng, Wenjun, et al.
Published: (2024)
by: Zeng, Wenjun, et al.
Published: (2024)
InstructAV: Instruction Fine-tuning Large Language Models for Authorship Verification
by: Hu, Yujia, et al.
Published: (2024)
by: Hu, Yujia, et al.
Published: (2024)
Language over Content: Tracing Cultural Understanding in Multilingual Large Language Models
by: Cho, Seungho, et al.
Published: (2025)
by: Cho, Seungho, et al.
Published: (2025)
Scaling Up LLM Reviews for Google Ads Content Moderation
by: Qiao, Wei, et al.
Published: (2024)
by: Qiao, Wei, et al.
Published: (2024)
Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models
by: Chua, Lynn, et al.
Published: (2024)
by: Chua, Lynn, et al.
Published: (2024)
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
Building a Few-Shot Cross-Domain Multilingual NLU Model for Customer Care
by: Kumar, Saurabh, et al.
Published: (2025)
by: Kumar, Saurabh, et al.
Published: (2025)
Building Multilingual Datasets for Predicting Mental Health Severity through LLMs: Prospects and Challenges
by: Skianis, Konstantinos, et al.
Published: (2024)
by: Skianis, Konstantinos, et al.
Published: (2024)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
by: Dinuta, Eduard Stefan, et al.
Published: (2025)
by: Dinuta, Eduard Stefan, et al.
Published: (2025)
On the Calibration of Multilingual Question Answering LLMs
by: Yang, Yahan, et al.
Published: (2023)
by: Yang, Yahan, et al.
Published: (2023)
Guarding the Meaning: Self-Supervised Training for Semantic Robustness in Guard Models
by: Pinneri, Cristina, et al.
Published: (2025)
by: Pinneri, Cristina, et al.
Published: (2025)
How Far Can 100 Samples Go? Unlocking Overall Zero-Shot Multilingual Translation via Tiny Multi-Parallel Data
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
NoisyHate: Mining Online Human-Written Perturbations for Realistic Robustness Benchmarking of Content Moderation Models
by: Ye, Yiran, et al.
Published: (2023)
by: Ye, Yiran, et al.
Published: (2023)
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
by: Ghosh, Shaona, et al.
Published: (2024)
by: Ghosh, Shaona, et al.
Published: (2024)
Let Community Rules Be Reflected in Online Content Moderation
by: Xin, Wangjiaxuan, et al.
Published: (2024)
by: Xin, Wangjiaxuan, et al.
Published: (2024)
FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation
by: Ding, Zhihao, et al.
Published: (2026)
by: Ding, Zhihao, et al.
Published: (2026)
D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning
by: Zhang, Ru, et al.
Published: (2026)
by: Zhang, Ru, et al.
Published: (2026)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
Enhancing Multilingual Sentiment Analysis with Explainability for Sinhala, English, and Code-Mixed Content
by: Rizvi, Azmarah, et al.
Published: (2025)
by: Rizvi, Azmarah, et al.
Published: (2025)
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
by: Wang, Yaxuan, et al.
Published: (2025)
by: Wang, Yaxuan, et al.
Published: (2025)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
by: Nagaraj, Manish, et al.
Published: (2025)
by: Nagaraj, Manish, et al.
Published: (2025)
Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey
by: Lee, Kunil, et al.
Published: (2026)
by: Lee, Kunil, et al.
Published: (2026)
Analysing Public Transport User Sentiment on Low Resource Multilingual Data
by: Myoya, Rozina L., et al.
Published: (2024)
by: Myoya, Rozina L., et al.
Published: (2024)
DynaGuard: A Dynamic Guardian Model With User-Defined Policies
by: Hoover, Monte, et al.
Published: (2025)
by: Hoover, Monte, et al.
Published: (2025)
CASE: Efficient Curricular Data Pre-training for Building Assistive Psychology Expert Models
by: Harne, Sarthak, et al.
Published: (2024)
by: Harne, Sarthak, et al.
Published: (2024)
TableGuard -- Securing Structured & Unstructured Data
by: Sharma, Anantha, et al.
Published: (2024)
by: Sharma, Anantha, et al.
Published: (2024)
Guided Distant Supervision for Multilingual Relation Extraction Data: Adapting to a New Language
by: Plum, Alistair, et al.
Published: (2024)
by: Plum, Alistair, et al.
Published: (2024)
WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
by: Li, Jiacheng, et al.
Published: (2025)
by: Li, Jiacheng, et al.
Published: (2025)
Massive Editing for Large Language Models via Meta Learning
by: Tan, Chenmien, et al.
Published: (2023)
by: Tan, Chenmien, et al.
Published: (2023)
Similar Items
-
Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
by: Chua, Gabriel, et al.
Published: (2025) -
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024) -
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
by: Ge, Ziyu, et al.
Published: (2025) -
On Calibration of LLM-based Guard Models for Reliable Content Moderation
by: Liu, Hongfu, et al.
Published: (2024) -
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
by: Elesedy, Hayder, et al.
Published: (2024)