Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
Fuente:
arXiv
Salvato in:
| Autori principali: | Chua, Gabriel, Tan, Leanne, Ge, Ziyu, Lee, Roy Ka-Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
di: Tan, Leanne, et al.
Pubblicazione: (2025)
di: Tan, Leanne, et al.
Pubblicazione: (2025)
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
di: Ge, Ziyu, et al.
Pubblicazione: (2025)
di: Ge, Ziyu, et al.
Pubblicazione: (2025)
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
di: Hwang, Jaedong, et al.
Pubblicazione: (2025)
di: Hwang, Jaedong, et al.
Pubblicazione: (2025)
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
di: Luo, Zheng, et al.
Pubblicazione: (2026)
di: Luo, Zheng, et al.
Pubblicazione: (2026)
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
di: Lee, Juyong, et al.
Pubblicazione: (2024)
di: Lee, Juyong, et al.
Pubblicazione: (2024)
MPO: Multilingual Safety Alignment via Reward Gap Optimization
di: Zhao, Weixiang, et al.
Pubblicazione: (2025)
di: Zhao, Weixiang, et al.
Pubblicazione: (2025)
Contrastive Token-level Explanations for Graph-based Rumour Detection
di: Chin, Daniel Wai Kit, et al.
Pubblicazione: (2025)
di: Chin, Daniel Wai Kit, et al.
Pubblicazione: (2025)
Multilingual Safety Alignment via Self-Distillation
di: Qin, Ruiyang, et al.
Pubblicazione: (2026)
di: Qin, Ruiyang, et al.
Pubblicazione: (2026)
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
di: Li, Junjie, et al.
Pubblicazione: (2026)
di: Li, Junjie, et al.
Pubblicazione: (2026)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
di: Kang, Deokhyung, et al.
Pubblicazione: (2025)
di: Kang, Deokhyung, et al.
Pubblicazione: (2025)
InstructAV: Instruction Fine-tuning Large Language Models for Authorship Verification
di: Hu, Yujia, et al.
Pubblicazione: (2024)
di: Hu, Yujia, et al.
Pubblicazione: (2024)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
di: Zhou, Yujun, et al.
Pubblicazione: (2024)
di: Zhou, Yujun, et al.
Pubblicazione: (2024)
Safe at the Margins: A General Approach to Safety Alignment in Low-Resource English Languages -- A Singlish Case Study
di: Lim, Isaac, et al.
Pubblicazione: (2025)
di: Lim, Isaac, et al.
Pubblicazione: (2025)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
di: Zheng, Weihua, et al.
Pubblicazione: (2026)
di: Zheng, Weihua, et al.
Pubblicazione: (2026)
Resolving Conflicting Evidence in Automated Fact-Checking: A Study on Retrieval-Augmented LLMs
di: Ge, Ziyu, et al.
Pubblicazione: (2025)
di: Ge, Ziyu, et al.
Pubblicazione: (2025)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
di: Gureja, Srishti, et al.
Pubblicazione: (2024)
di: Gureja, Srishti, et al.
Pubblicazione: (2024)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
di: Bu, Yuyan, et al.
Pubblicazione: (2026)
di: Bu, Yuyan, et al.
Pubblicazione: (2026)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
di: Gringras, David
Pubblicazione: (2026)
di: Gringras, David
Pubblicazione: (2026)
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
di: Wu, Yuhao, et al.
Pubblicazione: (2025)
di: Wu, Yuhao, et al.
Pubblicazione: (2025)
Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models
di: Chua, Lynn, et al.
Pubblicazione: (2024)
di: Chua, Lynn, et al.
Pubblicazione: (2024)
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
di: Hu, Yujia, et al.
Pubblicazione: (2025)
di: Hu, Yujia, et al.
Pubblicazione: (2025)
MOSLD-Bench: Multilingual Open-Set Learning and Discovery Benchmark for Text Categorization
di: Costache, Adriana-Valentina, et al.
Pubblicazione: (2026)
di: Costache, Adriana-Valentina, et al.
Pubblicazione: (2026)
Lost in the Prompt Order: Revealing the Limitations of Causal Attention in Language Models
di: Ok, Hyunjong, et al.
Pubblicazione: (2026)
di: Ok, Hyunjong, et al.
Pubblicazione: (2026)
Building a Few-Shot Cross-Domain Multilingual NLU Model for Customer Care
di: Kumar, Saurabh, et al.
Pubblicazione: (2025)
di: Kumar, Saurabh, et al.
Pubblicazione: (2025)
Building Multilingual Datasets for Predicting Mental Health Severity through LLMs: Prospects and Challenges
di: Skianis, Konstantinos, et al.
Pubblicazione: (2024)
di: Skianis, Konstantinos, et al.
Pubblicazione: (2024)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
di: Kumar, Somnath, et al.
Pubblicazione: (2024)
di: Kumar, Somnath, et al.
Pubblicazione: (2024)
Building from Scratch: A Multi-Agent Framework with Human-in-the-Loop for Multilingual Legal Terminology Mapping
di: Meng, Lingyi, et al.
Pubblicazione: (2025)
di: Meng, Lingyi, et al.
Pubblicazione: (2025)
QueryBuilder: Human-in-the-Loop Query Development for Information Retrieval
di: Kandula, Hemanth, et al.
Pubblicazione: (2024)
di: Kandula, Hemanth, et al.
Pubblicazione: (2024)
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
The Proxy Presumption: From Semantic Embeddings to Valid Social Measures
di: Li, Baishi, et al.
Pubblicazione: (2026)
di: Li, Baishi, et al.
Pubblicazione: (2026)
CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
di: Joshi, Raviraj, et al.
Pubblicazione: (2025)
di: Joshi, Raviraj, et al.
Pubblicazione: (2025)
Interpreting Bias in Large Language Models: A Feature-Based Approach
di: Prakash, Nirmalendu, et al.
Pubblicazione: (2024)
di: Prakash, Nirmalendu, et al.
Pubblicazione: (2024)
CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation
di: Chen, Tong, et al.
Pubblicazione: (2024)
di: Chen, Tong, et al.
Pubblicazione: (2024)
On the Calibration of Multilingual Question Answering LLMs
di: Yang, Yahan, et al.
Pubblicazione: (2023)
di: Yang, Yahan, et al.
Pubblicazione: (2023)
Pretrained Multilingual Transformers Reveal Quantitative Distance Between Human Languages
di: Zhao, Yue, et al.
Pubblicazione: (2026)
di: Zhao, Yue, et al.
Pubblicazione: (2026)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
di: Jin, Chang, et al.
Pubblicazione: (2026)
di: Jin, Chang, et al.
Pubblicazione: (2026)
MinorBench: A hand-built benchmark for content-based risks for children
di: Khoo, Shaun, et al.
Pubblicazione: (2025)
di: Khoo, Shaun, et al.
Pubblicazione: (2025)
Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data
di: Kwak, Minseo, et al.
Pubblicazione: (2026)
di: Kwak, Minseo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
di: Tan, Leanne, et al.
Pubblicazione: (2025) -
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
di: Ge, Ziyu, et al.
Pubblicazione: (2025) -
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
di: Hwang, Jaedong, et al.
Pubblicazione: (2025) -
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
di: Luo, Zheng, et al.
Pubblicazione: (2026) -
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
di: Lee, Juyong, et al.
Pubblicazione: (2024)