Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
Fuente:
arXiv
Saved in:
| Main Authors: | Chua, Gabriel, Tan, Leanne, Ge, Ziyu, Lee, Roy Ka-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
by: Tan, Leanne, et al.
Published: (2025)
by: Tan, Leanne, et al.
Published: (2025)
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
by: Ge, Ziyu, et al.
Published: (2025)
by: Ge, Ziyu, et al.
Published: (2025)
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
by: Hwang, Jaedong, et al.
Published: (2025)
by: Hwang, Jaedong, et al.
Published: (2025)
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
by: Luo, Zheng, et al.
Published: (2026)
by: Luo, Zheng, et al.
Published: (2026)
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
by: Lee, Juyong, et al.
Published: (2024)
by: Lee, Juyong, et al.
Published: (2024)
MPO: Multilingual Safety Alignment via Reward Gap Optimization
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Contrastive Token-level Explanations for Graph-based Rumour Detection
by: Chin, Daniel Wai Kit, et al.
Published: (2025)
by: Chin, Daniel Wai Kit, et al.
Published: (2025)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
by: Li, Junjie, et al.
Published: (2026)
by: Li, Junjie, et al.
Published: (2026)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
by: Kang, Deokhyung, et al.
Published: (2025)
by: Kang, Deokhyung, et al.
Published: (2025)
InstructAV: Instruction Fine-tuning Large Language Models for Authorship Verification
by: Hu, Yujia, et al.
Published: (2024)
by: Hu, Yujia, et al.
Published: (2024)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
Safe at the Margins: A General Approach to Safety Alignment in Low-Resource English Languages -- A Singlish Case Study
by: Lim, Isaac, et al.
Published: (2025)
by: Lim, Isaac, et al.
Published: (2025)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
by: Zheng, Weihua, et al.
Published: (2026)
by: Zheng, Weihua, et al.
Published: (2026)
Resolving Conflicting Evidence in Automated Fact-Checking: A Study on Retrieval-Augmented LLMs
by: Ge, Ziyu, et al.
Published: (2025)
by: Ge, Ziyu, et al.
Published: (2025)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
by: Gureja, Srishti, et al.
Published: (2024)
by: Gureja, Srishti, et al.
Published: (2024)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
by: Bu, Yuyan, et al.
Published: (2026)
by: Bu, Yuyan, et al.
Published: (2026)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models
by: Chua, Lynn, et al.
Published: (2024)
by: Chua, Lynn, et al.
Published: (2024)
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
by: Hu, Yujia, et al.
Published: (2025)
by: Hu, Yujia, et al.
Published: (2025)
MOSLD-Bench: Multilingual Open-Set Learning and Discovery Benchmark for Text Categorization
by: Costache, Adriana-Valentina, et al.
Published: (2026)
by: Costache, Adriana-Valentina, et al.
Published: (2026)
Lost in the Prompt Order: Revealing the Limitations of Causal Attention in Language Models
by: Ok, Hyunjong, et al.
Published: (2026)
by: Ok, Hyunjong, et al.
Published: (2026)
Building a Few-Shot Cross-Domain Multilingual NLU Model for Customer Care
by: Kumar, Saurabh, et al.
Published: (2025)
by: Kumar, Saurabh, et al.
Published: (2025)
Building Multilingual Datasets for Predicting Mental Health Severity through LLMs: Prospects and Challenges
by: Skianis, Konstantinos, et al.
Published: (2024)
by: Skianis, Konstantinos, et al.
Published: (2024)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
by: Kumar, Somnath, et al.
Published: (2024)
by: Kumar, Somnath, et al.
Published: (2024)
Building from Scratch: A Multi-Agent Framework with Human-in-the-Loop for Multilingual Legal Terminology Mapping
by: Meng, Lingyi, et al.
Published: (2025)
by: Meng, Lingyi, et al.
Published: (2025)
QueryBuilder: Human-in-the-Loop Query Development for Information Retrieval
by: Kandula, Hemanth, et al.
Published: (2024)
by: Kandula, Hemanth, et al.
Published: (2024)
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
The Proxy Presumption: From Semantic Embeddings to Valid Social Measures
by: Li, Baishi, et al.
Published: (2026)
by: Li, Baishi, et al.
Published: (2026)
CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
by: Joshi, Raviraj, et al.
Published: (2025)
by: Joshi, Raviraj, et al.
Published: (2025)
Interpreting Bias in Large Language Models: A Feature-Based Approach
by: Prakash, Nirmalendu, et al.
Published: (2024)
by: Prakash, Nirmalendu, et al.
Published: (2024)
CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
On the Calibration of Multilingual Question Answering LLMs
by: Yang, Yahan, et al.
Published: (2023)
by: Yang, Yahan, et al.
Published: (2023)
Pretrained Multilingual Transformers Reveal Quantitative Distance Between Human Languages
by: Zhao, Yue, et al.
Published: (2026)
by: Zhao, Yue, et al.
Published: (2026)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
by: Jin, Chang, et al.
Published: (2026)
by: Jin, Chang, et al.
Published: (2026)
MinorBench: A hand-built benchmark for content-based risks for children
by: Khoo, Shaun, et al.
Published: (2025)
by: Khoo, Shaun, et al.
Published: (2025)
Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data
by: Kwak, Minseo, et al.
Published: (2026)
by: Kwak, Minseo, et al.
Published: (2026)
Similar Items
-
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
by: Tan, Leanne, et al.
Published: (2025) -
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
by: Ge, Ziyu, et al.
Published: (2025) -
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
by: Hwang, Jaedong, et al.
Published: (2025) -
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
by: Luo, Zheng, et al.
Published: (2026) -
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
by: Lee, Juyong, et al.
Published: (2024)