Safe at the Margins: A General Approach to Safety Alignment in Low-Resource English Languages -- A Singlish Case Study
Fuente:
arXiv
Saved in:
| Main Authors: | Lim, Isaac, Khoo, Shaun, Lee, Roy Ka-Wei, Chua, Watson, Goh, Jia Yi, Foo, Jessica |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications
by: Goh, Jia Yi, et al.
Published: (2025)
by: Goh, Jia Yi, et al.
Published: (2025)
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
by: Ge, Ziyu, et al.
Published: (2025)
by: Ge, Ziyu, et al.
Published: (2025)
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
by: Khoo, Shaun, et al.
Published: (2025)
by: Khoo, Shaun, et al.
Published: (2025)
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024)
by: Foo, Jessica, et al.
Published: (2024)
Know Or Not: a library for evaluating out-of-knowledge base robustness
by: Foo, Jessica, et al.
Published: (2025)
by: Foo, Jessica, et al.
Published: (2025)
Disentangling Singlish Discourse Particles with Task-Driven Representation
by: Foo, Linus Tze En, et al.
Published: (2024)
by: Foo, Linus Tze En, et al.
Published: (2024)
MinorBench: A hand-built benchmark for content-based risks for children
by: Khoo, Shaun, et al.
Published: (2025)
by: Khoo, Shaun, et al.
Published: (2025)
Stylistic Evolution and LLM Neutrality in Singlish Language
by: Foo, Linus Tze En, et al.
Published: (2026)
by: Foo, Linus Tze En, et al.
Published: (2026)
From Standard English to Singlish: A Retrieval-Augmented Approach for Code-Switched Creole Generation in Large Language Models
by: Lai, Foong Ming, et al.
Published: (2026)
by: Lai, Foong Ming, et al.
Published: (2026)
A Flexible Large Language Models Guardrail Development Methodology Applied to Off-Topic Prompt Detection
by: Chua, Gabriel, et al.
Published: (2024)
by: Chua, Gabriel, et al.
Published: (2024)
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
by: Hu, Yujia, et al.
Published: (2025)
by: Hu, Yujia, et al.
Published: (2025)
Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
by: Chua, Gabriel, et al.
Published: (2025)
by: Chua, Gabriel, et al.
Published: (2025)
Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models
by: Tan, Rui Yang, et al.
Published: (2026)
by: Tan, Rui Yang, et al.
Published: (2026)
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
by: Wang, Zhaoxin, et al.
Published: (2026)
by: Wang, Zhaoxin, et al.
Published: (2026)
Facilitating acculturation of internationally educated nurses: A meta‐synthesis of social integration strategies
by: Huili Eugenia Foo, et al.
Published: (2024)
by: Huili Eugenia Foo, et al.
Published: (2024)
Interpreting Bias in Large Language Models: A Feature-Based Approach
by: Prakash, Nirmalendu, et al.
Published: (2024)
by: Prakash, Nirmalendu, et al.
Published: (2024)
Small Changes, Big Impact: Demographic Bias in LLM-Based Hiring Through Subtle Sociocultural Markers in Anonymised Resumes
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2026)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2026)
Limpeh ga li gong: Challenges in Singlish Annotations
by: Chan, Luo Qi, et al.
Published: (2024)
by: Chan, Luo Qi, et al.
Published: (2024)
Swa Bhasha: Message-Based Singlish to Sinhala Transliteration
by: Athukorala, Maneesha U., et al.
Published: (2024)
by: Athukorala, Maneesha U., et al.
Published: (2024)
SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia
by: Ng, Ri Chi, et al.
Published: (2026)
by: Ng, Ri Chi, et al.
Published: (2026)
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
by: Tan, Leanne, et al.
Published: (2025)
by: Tan, Leanne, et al.
Published: (2025)
Low-Resource NMT: A Case Study on the Written and Spoken Languages in Hong Kong
by: Mak, Hei Yi, et al.
Published: (2025)
by: Mak, Hei Yi, et al.
Published: (2025)
BGM-HAN: A Hierarchical Attention Network for Accurate and Fair Decision Assessment on Semi-Structured Profiles
by: Liu, Junhua, et al.
Published: (2025)
by: Liu, Junhua, et al.
Published: (2025)
Elderly Onset Primary Intestinal Lymphangiectasia—A Rare Case
by: Li‐Han Goh, et al.
Published: (2025)
by: Li‐Han Goh, et al.
Published: (2025)
SafeWorld: Geo-Diverse Safety Alignment
by: Yin, Da, et al.
Published: (2024)
by: Yin, Da, et al.
Published: (2024)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
Advancing LLM Safe Alignment with Safety Representation Ranking
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Measuring Safety Alignment Effects in Autonomous Security Agents
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases
by: Feakins, Shaun, et al.
Published: (2026)
by: Feakins, Shaun, et al.
Published: (2026)
Towards Objective and Unbiased Decision Assessments with LLM-Enhanced Hierarchical Attention Networks
by: Liu, Junhua, et al.
Published: (2024)
by: Liu, Junhua, et al.
Published: (2024)
Understanding Fairness-Accuracy Trade-offs in Machine Learning Models: Does Promoting Fairness Undermine Performance?
by: Liu, Junhua, et al.
Published: (2024)
by: Liu, Junhua, et al.
Published: (2024)
SGHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Singapore
by: Ng, Ri Chi, et al.
Published: (2024)
by: Ng, Ri Chi, et al.
Published: (2024)
Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
by: Ao, Shuang, et al.
Published: (2025)
by: Ao, Shuang, et al.
Published: (2025)
An Input-to-State Safety Approach Towards Safe Control of a Class of Parabolic PDEs Under Disturbances
by: Roy, Tanushree, et al.
Published: (2022)
by: Roy, Tanushree, et al.
Published: (2022)
Sinhala-English Word Embedding Alignment: Introducing Datasets and Benchmark for a Low Resource Language
by: Wickramasinghe, Kasun, et al.
Published: (2023)
by: Wickramasinghe, Kasun, et al.
Published: (2023)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
by: Feng, Jingyuan, et al.
Published: (2026)
by: Feng, Jingyuan, et al.
Published: (2026)
Three Approaches to the Automation of Laser System Alignment and Their Resource Implications: A Case Study
by: Robb, David A., et al.
Published: (2024)
by: Robb, David A., et al.
Published: (2024)
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Similar Items
-
Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications
by: Goh, Jia Yi, et al.
Published: (2025) -
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
by: Ge, Ziyu, et al.
Published: (2025) -
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
by: Khoo, Shaun, et al.
Published: (2025) -
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024) -
Know Or Not: a library for evaluating out-of-knowledge base robustness
by: Foo, Jessica, et al.
Published: (2025)