Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goh, Jia Yi, Khoo, Shaun, Iskandar, Nyx, Chua, Gabriel, Tan, Leanne, Foo, Jessica
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912480884686848
author Goh, Jia Yi
Khoo, Shaun
Iskandar, Nyx
Chua, Gabriel
Tan, Leanne
Foo, Jessica
author_facet Goh, Jia Yi
Khoo, Shaun
Iskandar, Nyx
Chua, Gabriel
Tan, Leanne
Foo, Jessica
contents Most safety testing efforts for large language models (LLMs) today focus on evaluating foundation models. However, there is a growing need to evaluate safety at the application level, as components such as system prompts, retrieval pipelines, and guardrails introduce additional factors that significantly influence the overall safety of LLM applications. In this paper, we introduce a practical framework for evaluating application-level safety in LLM systems, validated through real-world deployment across multiple use cases within our organization. The framework consists of two parts: (1) principles for developing customized safety risk taxonomies, and (2) practices for evaluating safety risks in LLM applications. We illustrate how the proposed framework was applied in our internal pilot, providing a reference point for organizations seeking to scale their safety testing efforts. This work aims to bridge the gap between theoretical concepts in AI safety and the operational realities of safeguarding LLM applications in practice, offering actionable guidance for safe and scalable deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2507_09820
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications
Goh, Jia Yi
Khoo, Shaun
Iskandar, Nyx
Chua, Gabriel
Tan, Leanne
Foo, Jessica
Software Engineering
Computers and Society
Most safety testing efforts for large language models (LLMs) today focus on evaluating foundation models. However, there is a growing need to evaluate safety at the application level, as components such as system prompts, retrieval pipelines, and guardrails introduce additional factors that significantly influence the overall safety of LLM applications. In this paper, we introduce a practical framework for evaluating application-level safety in LLM systems, validated through real-world deployment across multiple use cases within our organization. The framework consists of two parts: (1) principles for developing customized safety risk taxonomies, and (2) practices for evaluating safety risks in LLM applications. We illustrate how the proposed framework was applied in our internal pilot, providing a reference point for organizations seeking to scale their safety testing efforts. This work aims to bridge the gap between theoretical concepts in AI safety and the operational realities of safeguarding LLM applications in practice, offering actionable guidance for safe and scalable deployment.
title Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications
topic Software Engineering
Computers and Society
url https://arxiv.org/abs/2507.09820