LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915518627184640 |
|---|---|
| author | Tan, Leanne Chua, Gabriel Ge, Ziyu Lee, Roy Ka-Wei |
| author_facet | Tan, Leanne Chua, Gabriel Ge, Ziyu Lee, Roy Ka-Wei |
| contents | Modern moderation systems increasingly support multiple languages, but often fail to address localisation and low-resource variants - creating safety gaps in real-world deployments. Small models offer a potential alternative to large LLMs, yet still demand considerable data and compute. We present LionGuard 2, a lightweight, multilingual moderation classifier tailored to the Singapore context, supporting English, Chinese, Malay, and partial Tamil. Built on pre-trained OpenAI embeddings and a multi-head ordinal classifier, LionGuard 2 outperforms several commercial and open-source systems across 17 benchmarks, including both Singapore-specific and public English datasets. The system is actively deployed within the Singapore Government, demonstrating practical efficacy at scale. Our findings show that high-quality local data and robust multilingual embeddings can achieve strong moderation performance, without fine-tuning large models. We release our model weights and part of our training data to support future work on LLM safety. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_15339 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators Tan, Leanne Chua, Gabriel Ge, Ziyu Lee, Roy Ka-Wei Computation and Language Machine Learning Modern moderation systems increasingly support multiple languages, but often fail to address localisation and low-resource variants - creating safety gaps in real-world deployments. Small models offer a potential alternative to large LLMs, yet still demand considerable data and compute. We present LionGuard 2, a lightweight, multilingual moderation classifier tailored to the Singapore context, supporting English, Chinese, Malay, and partial Tamil. Built on pre-trained OpenAI embeddings and a multi-head ordinal classifier, LionGuard 2 outperforms several commercial and open-source systems across 17 benchmarks, including both Singapore-specific and public English datasets. The system is actively deployed within the Singapore Government, demonstrating practical efficacy at scale. Our findings show that high-quality local data and robust multilingual embeddings can achieve strong moderation performance, without fine-tuning large models. We release our model weights and part of our training data to support future work on LLM safety. |
| title | LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2507.15339 |