Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Goh, Jia Yi, Khoo, Shaun, Iskandar, Nyx, Chua, Gabriel, Tan, Leanne, Foo, Jessica |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safe at the Margins: A General Approach to Safety Alignment in Low-Resource English Languages -- A Singlish Case Study
by: Lim, Isaac, et al.
Published: (2025)
by: Lim, Isaac, et al.
Published: (2025)
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024)
by: Foo, Jessica, et al.
Published: (2024)
A Matter of Representation: Towards Graph-Based Abstract Code Generation
by: Iskandar, Nyx, et al.
Published: (2025)
by: Iskandar, Nyx, et al.
Published: (2025)
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
by: Khoo, Shaun, et al.
Published: (2025)
by: Khoo, Shaun, et al.
Published: (2025)
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
by: Tan, Honghao, et al.
Published: (2026)
by: Tan, Honghao, et al.
Published: (2026)
Knowledge Lever Risk Management for Software Engineering: A Stochastic Framework for Mitigating Knowledge Loss
by: Chua, Mark, et al.
Published: (2026)
by: Chua, Mark, et al.
Published: (2026)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
by: Chi, Wayne, et al.
Published: (2025)
by: Chi, Wayne, et al.
Published: (2025)
A Taxonomy of Real-World Defeaters in Safety Assurance Cases
by: Gohar, Usman, et al.
Published: (2025)
by: Gohar, Usman, et al.
Published: (2025)
Shedding Light onto Safety Integrity Level and Basic Software Constraints in a Real-World Automotive Application: Case Study with Driverator Framework
by: Denzinger, Tobias, et al.
Published: (2026)
by: Denzinger, Tobias, et al.
Published: (2026)
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
CIRCLE: A Framework for Evaluating AI from a Real-World Lens
by: Schwartz, Reva, et al.
Published: (2026)
by: Schwartz, Reva, et al.
Published: (2026)
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026)
by: Deng, Gelei, et al.
Published: (2026)
Know Or Not: a library for evaluating out-of-knowledge base robustness
by: Foo, Jessica, et al.
Published: (2025)
by: Foo, Jessica, et al.
Published: (2025)
AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Evaluating LLMs Code Reasoning Under Real-World Context
by: Liu, Changshu
Published: (2026)
by: Liu, Changshu
Published: (2026)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
by: Agarwal, Anisha, et al.
Published: (2024)
by: Agarwal, Anisha, et al.
Published: (2024)
Testing Autonomous Driving Systems -- What Really Matters and What Doesn't
by: Li, Changwen, et al.
Published: (2025)
by: Li, Changwen, et al.
Published: (2025)
MinorBench: A hand-built benchmark for content-based risks for children
by: Khoo, Shaun, et al.
Published: (2025)
by: Khoo, Shaun, et al.
Published: (2025)
TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
by: Jain, Kush, et al.
Published: (2024)
by: Jain, Kush, et al.
Published: (2024)
A Fuzzy Approach to Project Success: Measuring What Matters
by: Granja-Correia, João, et al.
Published: (2025)
by: Granja-Correia, João, et al.
Published: (2025)
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
by: Li, Lehui, et al.
Published: (2026)
by: Li, Lehui, et al.
Published: (2026)
An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains
by: Yan, Zihe, et al.
Published: (2025)
by: Yan, Zihe, et al.
Published: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
by: Xu, Yisen, et al.
Published: (2026)
by: Xu, Yisen, et al.
Published: (2026)
Reconciling Safety Measurement and Dynamic Assurance
by: Denney, Ewen, et al.
Published: (2024)
by: Denney, Ewen, et al.
Published: (2024)
Do AI Models Dream of Faster Code? An Empirical Study on LLM-Proposed Performance Improvements in Real-World Software
by: Yi, Lirong, et al.
Published: (2025)
by: Yi, Lirong, et al.
Published: (2025)
Evaluating Code Reasoning Abilities of Large Language Models Under Real-World Settings
by: Liu, Changshu, et al.
Published: (2025)
by: Liu, Changshu, et al.
Published: (2025)
Belobog: Move Language Fuzzing Framework For Real-World Smart Contracts
by: Kong, Ziqiao, et al.
Published: (2025)
by: Kong, Ziqiao, et al.
Published: (2025)
SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated Code
by: Wang, Qinglin, et al.
Published: (2025)
by: Wang, Qinglin, et al.
Published: (2025)
Should Code Models Learn Pedagogically? A Preliminary Evaluation of Curriculum Learning for Real-World Software Engineering Tasks
by: Khant, Kyi Shin, et al.
Published: (2025)
by: Khant, Kyi Shin, et al.
Published: (2025)
ContractTinker: LLM-Empowered Vulnerability Repair for Real-World Smart Contracts
by: Wang, Che, et al.
Published: (2024)
by: Wang, Che, et al.
Published: (2024)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
A Benchmark for Language Models in Real-World System Building
by: Jin, Weilin, et al.
Published: (2026)
by: Jin, Weilin, et al.
Published: (2026)
CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
by: Hu, Li, et al.
Published: (2025)
by: Hu, Li, et al.
Published: (2025)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects
by: Nunes, Henrique, et al.
Published: (2025)
by: Nunes, Henrique, et al.
Published: (2025)
A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models
by: Hu, Ruida, et al.
Published: (2024)
by: Hu, Ruida, et al.
Published: (2024)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
Decompile-Bench: Million-Scale Binary-Source Function Pairs for Real-World Binary Decompilation
by: Tan, Hanzhuo, et al.
Published: (2025)
by: Tan, Hanzhuo, et al.
Published: (2025)
DistillSeq: A Framework for Safety Alignment Testing in Large Language Models using Knowledge Distillation
by: Yang, Mingke, et al.
Published: (2024)
by: Yang, Mingke, et al.
Published: (2024)
Similar Items
-
Safe at the Margins: A General Approach to Safety Alignment in Low-Resource English Languages -- A Singlish Case Study
by: Lim, Isaac, et al.
Published: (2025) -
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024) -
A Matter of Representation: Towards Graph-Based Abstract Code Generation
by: Iskandar, Nyx, et al.
Published: (2025) -
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
by: Khoo, Shaun, et al.
Published: (2025) -
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
by: Tan, Honghao, et al.
Published: (2026)