Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Seongmin, Cho, Aeree, Kim, Grace C., Peng, ShengYun, Phute, Mansi, Chau, Duen Horng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shape it Up! Restoring LLM Safety during Finetuning
by: Peng, ShengYun, et al.
Published: (2025)
by: Peng, ShengYun, et al.
Published: (2025)
LLM Attributor: Interactive Visual Attribution for LLM Generation
by: Lee, Seongmin, et al.
Published: (2024)
by: Lee, Seongmin, et al.
Published: (2024)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
by: Phute, Mansi, et al.
Published: (2023)
by: Phute, Mansi, et al.
Published: (2023)
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
by: Peng, ShengYun, et al.
Published: (2024)
by: Peng, ShengYun, et al.
Published: (2024)
Self-Supervised Pre-Training for Table Structure Recognition Transformer
by: Peng, ShengYun, et al.
Published: (2024)
by: Peng, ShengYun, et al.
Published: (2024)
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
by: Tan, Honghao, et al.
Published: (2026)
by: Tan, Honghao, et al.
Published: (2026)
UniTable: Towards a Unified Framework for Table Recognition via Self-Supervised Pretraining
by: Peng, ShengYun, et al.
Published: (2024)
by: Peng, ShengYun, et al.
Published: (2024)
Towards Bridging Formal Methods and Human Interpretability
by: Paul, Abhijit, et al.
Published: (2025)
by: Paul, Abhijit, et al.
Published: (2025)
Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
by: Cartagena, Arnold, et al.
Published: (2026)
by: Cartagena, Arnold, et al.
Published: (2026)
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
by: Yu, Shasha, et al.
Published: (2026)
by: Yu, Shasha, et al.
Published: (2026)
UNDREAM: Bridging Differentiable Rendering and Photorealistic Simulation for End-to-end Adversarial Attacks
by: Phute, Mansi, et al.
Published: (2025)
by: Phute, Mansi, et al.
Published: (2025)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
Can LLM Generate Regression Tests for Software Commits?
by: Liu, Jing, et al.
Published: (2025)
by: Liu, Jing, et al.
Published: (2025)
Event-Chain Analysis for Automated Driving and ADAS Systems: Ensuring Safety and Meeting Regulatory Timing Requirements
by: Dingler, Sebastian, et al.
Published: (2025)
by: Dingler, Sebastian, et al.
Published: (2025)
Testing Research Software: An In-Depth Survey of Practices, Methods, and Tools
by: Eisty, Nasir U., et al.
Published: (2025)
by: Eisty, Nasir U., et al.
Published: (2025)
Refining Fuzzed Crashing Inputs for Better Fault Diagnosis
by: Kim, Kieun, et al.
Published: (2025)
by: Kim, Kieun, et al.
Published: (2025)
SafePlanner: Testing Safety of the Automated Driving System Plan Model
by: Kim, Dohyun, et al.
Published: (2026)
by: Kim, Dohyun, et al.
Published: (2026)
Explaining Code Risk in OSS: Towards LLM-Generated Fault Prediction Interpretations
by: Adejumo, Elijah Kayode, et al.
Published: (2025)
by: Adejumo, Elijah Kayode, et al.
Published: (2025)
iCodeReviewer: Improving Secure Code Review with Mixture of Prompts
by: Peng, Yun, et al.
Published: (2025)
by: Peng, Yun, et al.
Published: (2025)
Cottontail: Large Language Model-Driven Concolic Execution for Highly Structured Test Input Generation
by: Tu, Haoxin, et al.
Published: (2025)
by: Tu, Haoxin, et al.
Published: (2025)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025)
by: Xia, Hongfei, et al.
Published: (2025)
Developing Compelling Safety Cases
by: Hawkins, Richard
Published: (2025)
by: Hawkins, Richard
Published: (2025)
Safety Factories - a Manifesto
by: Cârlan, Carmen, et al.
Published: (2025)
by: Cârlan, Carmen, et al.
Published: (2025)
Usability as a Weapon: Attacking the Safety of LLM-Based Code Generation via Usability Requirements
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
A Structured Approach to Safety Case Construction for AI Systems
by: Lee, Sung Une, et al.
Published: (2026)
by: Lee, Sung Une, et al.
Published: (2026)
An Analysis of Early-Stage Functional Safety Analysis Methods and Their Integration into Model-Based Systems Engineering
by: Shefa, Jannatul, et al.
Published: (2025)
by: Shefa, Jannatul, et al.
Published: (2025)
An LLM-driven Scenario Generation Pipeline Using an Extended Scenic DSL for Autonomous Driving Safety Validation
by: Safa, Fida Khandaker, et al.
Published: (2026)
by: Safa, Fida Khandaker, et al.
Published: (2026)
Reconciling Safety Measurement and Dynamic Assurance
by: Denney, Ewen, et al.
Published: (2024)
by: Denney, Ewen, et al.
Published: (2024)
LLM-Empowered Functional Safety and Security by Design in Automotive Systems
by: Petrovic, Nenad, et al.
Published: (2026)
by: Petrovic, Nenad, et al.
Published: (2026)
ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Enhancing LLM-Based Coding Tools through Native Integration of IDE-Derived Static Context
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
A Survey on the Techniques and Tools for Automated Requirements Elicitation and Analysis of Mobile Apps
by: Wang, Chong, et al.
Published: (2025)
by: Wang, Chong, et al.
Published: (2025)
Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
by: Jin, Can, et al.
Published: (2026)
by: Jin, Can, et al.
Published: (2026)
Safety Verification and Optimization in Industrial Drive Systems
by: Hasrat, Imran Riaz, et al.
Published: (2025)
by: Hasrat, Imran Riaz, et al.
Published: (2025)
Nested Fusion: A Method for Learning High Resolution Latent Structure of Multi-Scale Measurement Data on Mars
by: Wright, Austin P., et al.
Published: (2024)
by: Wright, Austin P., et al.
Published: (2024)
B-OCL: An Object Constraint Language Interpreter in Python
by: Haq, Fitash Ul, et al.
Published: (2025)
by: Haq, Fitash Ul, et al.
Published: (2025)
ShellFuzzer: Grammar-based Fuzzing of Shell Interpreters
by: Felici, Riccardo, et al.
Published: (2024)
by: Felici, Riccardo, et al.
Published: (2024)
LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-based Code Completion
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
Causality-aware Safety Testing for Autonomous Driving Systems
by: Tang, Wenbing, et al.
Published: (2025)
by: Tang, Wenbing, et al.
Published: (2025)
Similar Items
-
Shape it Up! Restoring LLM Safety during Finetuning
by: Peng, ShengYun, et al.
Published: (2025) -
LLM Attributor: Interactive Visual Attribution for LLM Generation
by: Lee, Seongmin, et al.
Published: (2024) -
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
by: Phute, Mansi, et al.
Published: (2023) -
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
by: Peng, ShengYun, et al.
Published: (2024) -
Self-Supervised Pre-Training for Table Structure Recognition Transformer
by: Peng, ShengYun, et al.
Published: (2024)