Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhuowei, Zhang, Bowei, Lin, Nankai, Hou, Tian, Wang, Lianxi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911206258769920
author Chen, Zhuowei
Zhang, Bowei
Lin, Nankai
Hou, Tian
Wang, Lianxi
author_facet Chen, Zhuowei
Zhang, Bowei
Lin, Nankai
Hou, Tian
Wang, Lianxi
contents Recent advances in LLMs have enhanced AI capabilities, but also increased the risk posed by malicious requests, highlighting the need for effective LLM safeguards to detect such queries. Existing approaches largely rely on classifier-based methods that lack interpretability and perform poorly on low-resource languages. To address these limitations, we propose ConsistentGuard, a novel reasoning-based multilingual safeguard, which enhances explainability via reasoning and boosts knowledge transfer between languages through alignment. With only 1,000 training samples, our method demonstrates superior performance on three datasets across six languages, outperforming larger models trained with significantly more data, and exhibits strong interpretability and generalization ability. We also contribute a multilingual benchmark extension and release our codes to support future research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10677
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data
Chen, Zhuowei
Zhang, Bowei
Lin, Nankai
Hou, Tian
Wang, Lianxi
Computation and Language
Recent advances in LLMs have enhanced AI capabilities, but also increased the risk posed by malicious requests, highlighting the need for effective LLM safeguards to detect such queries. Existing approaches largely rely on classifier-based methods that lack interpretability and perform poorly on low-resource languages. To address these limitations, we propose ConsistentGuard, a novel reasoning-based multilingual safeguard, which enhances explainability via reasoning and boosts knowledge transfer between languages through alignment. With only 1,000 training samples, our method demonstrates superior performance on three datasets across six languages, outperforming larger models trained with significantly more data, and exhibits strong interpretability and generalization ability. We also contribute a multilingual benchmark extension and release our codes to support future research.
title Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data
topic Computation and Language
url https://arxiv.org/abs/2510.10677